πŸ“Š TradingView Backtest Deep Validation

Period 2019-12-15 – 2026-08-22 Β· 923 trades Β· 2,441 days (6.7 yrs)
Leverage 1Γ— Commission/side 0.04% (RT 0.08%) Slippage 0% Profit share 20% (per trade) Initial capital 10,000 USDT Risk-free rate 2%
Total trades
923
2019-12-15 – 2026-08-22 Β· 2,441 days
Cumulative return
8,950,853.6%
after fees
Annualized return
450.8%
138 trades/yr
Sharpe
2.92
1.5+ good / 2.0+ excellent
MDD (closed trades)
-24.61%
incl. intratrade lows -25.23%
Profit factor
2.05
1.5+ good
Win rate
43.2%
payoff 3.89

Equity Curve & Drawdown

10^510^610^710^8 2019202120222023202420252026 USDT
Strategy equity (month-end, log)
Longest drawdown
116 days
51 trades Β· 2023-07-14 ~ 2023-11-08
Days since last high
106 days
Drawdown episodes
105
peak β†’ full recovery
Avg drawdown (trough)
-4.32%
median -3.44%
Avg underwater depth
-5.40%
avg depth while in drawdown
Avg recovery time
17 days
median 13 days
The chart aggregates month-end values only. Max drawdown is -24.61% on a per-trade basis; intra-month troughs look milder on a month-end curve. Drawdown duration measures psychological pain β€” the longest stretch without a new equity high was 116 days.

Top-5 Longest Drawdowns (peak β†’ new peak)

StartNew peakDurationTradesMax depth
2023-07-142023-11-08116 days51-8.6%
2023-04-142023-07-1491 days47-15.4%
2020-05-062020-07-2680 days42-14.2%
2023-12-112024-01-2646 days29-8.9%
2025-12-142026-01-2945 days22-10.9%

Performance Metrics

MetricValueBenchmark
Sharpe Ratio2.921.5+ good Β· 2.0+ excellent
Sortino Ratio15.972.0+ good
Calmar Ratio18.312.0+ good Β· 3.0+ excellent
Profit Factor2.05gross profit 3,631,551,095 / gross loss 1,769,562,261 USDT
Recovery Factor5.02net profit 1,861,988,833 / MDD 370,559,835 USDT
Kelly fraction (full)28.64%half Kelly 14.32%
Expectancy per trade1.377%138 USDT/trade (on initial capital)
Avg win / avg loss4.81% / 1.24%payoff 3.89
MDD (closed trades)-24.61%within βˆ’25% is safe
MDD (incl. intratrade lows)-25.23%TradingView method
Risk of Ruin0.0000%Z=0.3374 Β· 40 losses to ruin

Risk & Distribution

MetricValueDescription
Skewness3.102positive = long profit tail
Excess kurtosis10.878higher = fatter tails
Bootstrap 95% CI[1.020%, 1.750%]1,000 resamples
Monte Carlo 95% / 99% worst MDD-31.97% / -37.52%bootstrap
Shuffle 95% / 99% worst MDD-30.82% / -36.63%reorder only
Monte Carlo profitable scenarios100%median 8,473,864% Β· range [865,429%, 107,127,486%]
Longest losing / winning streak9 / 8theoretical max streak 12.1
Capital hit at worst streakβˆ’10.6%loss unit 1.24%/trade

Losing-Streak Distribution

13579
X = streak length Β· Y = occurrences

MDD Distribution β€” Simulation Methods Compared

actual -24.6%-42%-34%-27%-19%-12%
Shuffle (reorder only)Bootstrap (resampling)
Merely reordering the trades made 28% of scenarios suffer a deeper MDD than the actual -24.6%. Design margin and psychological stop limits around the 99% value (-36.6%), not the realized MDD. There were 229 losing streaks in total, averaging 2.3 losses.

Random Missed-Trade Simulation (400 runs each) — alert→API fill risk

Miss rateMedian multipleWorst 5%Profit probability
5%51,634.1Γ—27,294.9Γ—100%
10%29,060.0Γ—12,880.1Γ—100%
20%8,598.0Γ—3,586.9Γ—100%

Statistical Tests

TestResultVerdict
T-test (vs random entry)T = 7.630 Β· p = 0.0000βœ… significant at 1%
Runs Testactual 458 Β· expected 454.0 Β· Z=0.00 Β· p=0.9998βœ… random
Deflated Sharpe Ratio100.0%βœ… significant
PBO (overfitting probability)0.0% Β· 0 of 20 splits lost OOSβœ… low
IS/OOS (70/30)IS 3.21 β†’ OOS 2.27 (retention 71%)βœ… good
Execution-delay sensitivity1-trade delay β†’ 2.96 (1.2% change)βœ… robust
Look-ahead diagnosticssame-bar 0 Β· min hold 20 min Β· MFE capture 69%βœ… no red flags

Deflated Sharpe Ratio #32

target 95% 100.0% 0255075100
Probability that the Sharpe exceeds the chance maximum after correcting for multiple-parameter search bias. Higher is better; 95%+ is statistically significant.

PBO β€” Overfitting Probability #33

target 25% 0.0% 0255075100
Share of splits where the in-sample winner loses out-of-sample. Lower is better; below 25% means low overfitting risk.

Random Entry Test (t-distribution) #25

critical 1.65T=7.63

Sharpe by Execution Delay #24

2.922.962.962.952.9501235

Performance by Market Regime

SegmentTradesMean returnWin rateSharpe
Early (1/3)3072.14%43.6%3.72
Middle (2/3)3070.93%42.7%2.47
Late (3/3)3091.07%43.4%2.53

Mean Return by Segment

2.14%0.93%1.07%Early (1/3)Middle (2/3)Late (3/3)

Historical Crisis Windows

Crisis windowExposedEntries in windowStrategy returnWindow MDDWin rateMarket over windowVerdict
COVID crash (2020-02–04)99+70.1%-5.4%56%-42.3%βœ… defended
May-2021 crash (2021-05–07)1918+19.7%-10.8%32%-34.8%βœ… defended
LUNA collapse (2022-05–06)1919+33.9%-10.6%47%-57.2%βœ… defended
FTX bankruptcy (2022-11–12)1413+23.7%-5.6%57%-25.4%βœ… defended
Yen-carry unwind (2024-08)1918+34.7%-6.8%47%-33.2%βœ… defended
Tariff shock (2025-02–04)2625+50.5%-6.1%50%-44.1%βœ… defended
Oct-2025 mass liquidation (10-10–11)10+17.4%0.0%100%-22.0%βœ… defended
Of the 7 crisis windows with open exposure, 7 were defended successfully. An exposed trade is one whose holding period (entry–exit) overlapped the crisis by at least a day; counting by exit time alone would mix in post-crisis trades, so overlap is used instead.

Halving-Cycle Breakdown

CycleTradesCompound returnSharpeMDDWin rate
Pre-3rd halving (–2020-05)39233.9%5.20-6.0%51%
3rd cycle (2020-05 – 2024-04)52694,082.7%3.05-15.4%43%
4th cycle (2024-04 –)3582,746.5%2.47-24.6%42%

Volatility-Regime Breakdown

πŸ”₯ High-volatility regime
709,501.3%
451 trades Β· Sharpe 3.76 Β· WR 46%
🧊 Low-volatility regime
1,161.4%
472 trades Β· Sharpe 1.96 Β· WR 41%
High/low volatility is split at the median of the 20-trade rolling standard deviation of price log-returns.

Yearly & Monthly Returns

Compound Return by Year

3%1,250%916%636%67%502%544%33%20192020202120222023202420252026

Monthly Return Heatmap (%)

12345678910111220193202017511654-8132632138187120212752-048301135814112920223716143151661214181011202330-1511-853110542024102414192314262025294202516306921-551241226184202612-1152-91-721

Time Patterns

Profit Sum by Exit Hour (simple %)

0246810121416182022

Profit Sum by Weekday (simple %)

27611778124129247300MonTueWedThuFriSatSun

Monthly Trade Frequency

2019-122020-032020-062020-092020-122021-032021-062021-092021-122022-032022-062022-092022-122023-032023-062023-092023-122024-032024-062024-092024-122025-032025-062025-092025-122026-032026-06
The strongest exit hour is 5:00 (+150%) and the weakest 3:00 (-8%), with the top-3 hours concentrating 27% of profit. By weekday, Sun is strongest and Wed weakest; the weekend share is 43%. Last-12-month average 13.2 trades/month vs overall 11.4 (+16%).

Period Win Rates & Volatility

Profitable days
49%
676 trading days
Profitable weeks
59%
291 weeks
Profitable months
90%
81 months
Monthly volatility (Οƒ)
17.0%
annualized 58.8%

Anti-Overfitting β€” Walk-Forward & K-Fold

WFE Β· rolling
83%
60%+ robust
WFE Β· anchored
74%
expanding window
Profitable OOS steps
6/6
Embargo buffer
18 trades
isolates train↔test
StepTest windowIS SharpeOOS SharpeRetentionOOS return
12021-03-06 ~ 2022-03-093.163.81121%778.3%
22022-04-10 ~ 2023-03-204.543.5979%461.7%
32023-05-07 ~ 2023-12-223.480.9527%18.8%
42024-01-21 ~ 2024-12-100.703.36482%441.9%
52025-01-12 ~ 2025-11-283.113.18103%426.0%
62026-01-01 ~ 2026-08-083.270.278%3.0%

Train vs Test Sharpe by Step

IS Sharpe (train)OOS Sharpe (test) 123456

Purged & Embargoed 6-Fold Cross-Validation

FoldTest windowTrain SharpeTest SharpeTest returnPurged
12019-12-15 ~ 2021-03-132.793.882,666.4%18
22021-03-15 ~ 2022-06-282.593.551,171.6%36
32022-07-04 ~ 2023-07-223.022.79350.5%36
42023-07-23 ~ 2024-08-193.072.16193.1%36
52024-08-21 ~ 2025-09-012.843.12754.2%36
62025-09-01 ~ 2026-08-223.151.81125.6%18

Train vs Test Sharpe by Fold

Train SharpeTest Sharpe 123456
6 of 6 folds were profitable in their test window. Walk-forward validates in time order while K-fold rotates every segment through validation β€” similar results from both are cross-evidence of low overfitting risk.

Leverage Analysis

Kelly-optimal leverage
11.5Γ—
maximizes compounding
Half Kelly (recommended)
5.8Γ—
half the volatility
1%-ruin safe cap
9.6Γ—
ruin = 50% capital loss
Continuous Kelly (ΞΌβˆ’rf)/σ²
4.5Γ—
Markowitz tangency solution
Ruin at current 1Γ—
0.00%
Sharpe 2.92

Long-Run Compound Growth by Leverage (Kelly)

max 11.5Γ— 2Γ—4Γ—6Γ—8Γ—10Γ—12Γ—14Γ—16Γ—18Γ—20Γ—CAGR %
LeverageSharpeMDDRuin prob.CAGR (%/yr)Cumulative return
0.5Γ—2.8913.0%0.00%145.7%40,600.8%
1.0Γ— πŸ“current2.9224.6%0.00%450.8%8,950,853.6%
2.0Γ—2.9444.1%0.00%2,108.4%96,063,059,758.4%
3.0Γ—2.9459.3%0.00%6,845.3%2.033e+14%
5.0Γ—2.9579.4%0.02%39,201.4%2.181e+19%
9.6Γ— πŸ›‘οΈsafe cap2.9596.6%1.02%280,239.2%1.099e+25%
10.0Γ—2.9597.2%1.23%298,719.3%1.685e+25%
11.5Γ— 🎯Kelly2.9598.6%2.18%331,605.8%3.385e+25%
The discrete scan puts Kelly at 11.5Γ— while the closed-form continuous solution gives 4.5Γ—. The wider the gap, the further the return distribution sits from a lognormal approximation β€” a common signature of a stop-loss that artificially truncates the left tail. The textbook approach is to start live at the lower of half Kelly (5.8Γ—) and the safe cap (9.6Γ—).
CAGR and cumulative return in the high-leverage range are theoretical. They assume full profit reinvestment, no margin limits, no funding fees, and fixed execution slippage β€” so they are not achievable targets. Weight the ruin-probability and MDD columns far more heavily.

🦒 Black-Swan Injection Stress (βˆ’30% at 1Γ—, injected 3Γ—)

MetricBeforeAfter
Kelly-optimal leverage11.5Γ—11.5Γ—
Half Kelly (recommended)5.8Γ—5.8Γ—
1%-ruin safe cap9.6Γ—7.5Γ—
95% worst MDD (current 1Γ—)-30.8%-54.4%
History only contains the crises that actually happened. Injecting an unseen βˆ’30% event roughly once every two years lowers the Kelly optimum from 11.5Γ— to 11.5Γ—. The more honest live ceiling is the post-injection figure β€” the lower of half Kelly 5.8Γ— and the safe cap 7.5Γ—.

Costs, Slippage & Funding

ScenarioFinal equity multipleCumulative returnSharpeMDD
β‘  Zero cost (ideal)186,541.23Γ—18,654,023.4%3.12-20.8%
β‘‘ Maker RT 0.04%129,227.16Γ—12,922,616.2%3.04-22.7%
β‘’ Taker RT 0.10%74,490.90Γ—7,448,990.3%2.91-25.5%
β‘£ Taker+slippage RT 0.16%42,924.94Γ—4,292,394.2%2.78-28.3%
β‘€ Current setting RT 0.08%89,509.54Γ—8,950,853.6%2.95-24.6%

Slippage Sensitivity

Slippage (per side)Round-tripFinal equity multipleCumulative returnSharpe
0bp0bp89,509.54Γ—8,950,853.6%2.95
1bp2bp74,490.90Γ—7,448,990.3%2.91
2bp4bp61,989.95Γ—6,198,895.1%2.87
5bp10bp35,717.42Γ—3,571,642.0%2.74
10bp20bp14,239.45Γ—1,423,844.6%2.52

Funding-Fee Approximation (perpetuals, 0.01%/8h)

Avg holding time
28.1 h
avg 337 bars Γ— 5 min
Total funding cost
32.45%
vs capital
Equity multiple after funding
65,838.32Γ—
before 89,509.54Γ—
Sharpe after funding
2.89
-0.06
An average hold of 28 hours means about 3.5 funding events per trade, and funding alone removes 26% of the final equity. At the current setting (0.04% per side), each trade costs roughly 0.0800% of capital, compounded 923 times.

Profit Structure Decomposition

Top-Trade Removal β€” Luck Dependence

TestTop-trade profit shareMultiple after removalSharpe after removalVerdict
Remove top 57.8%24,193.01Γ—2.83βœ… still profitable
Remove top 1015.6%6,538.98Γ—2.71βœ… still profitable
Remove top 2027.5%842.59Γ—2.42βœ… still profitable

Long / Short Breakdown

SideTradesWin rateCompound returnSharpeMDD
🟒 Long63542.2%104,702.2%2.61-20.7%
πŸ”΄ Short28845.5%8,440.8%3.75-16.7%
Long–short monthly corr ρ
-0.30
more negative = internal hedge
Monthly vol (long / short)
14.6% / 10.2%
Combined monthly vol
15.1%
if uncorrelated 17.8%
Long–short monthly correlation ρ = -0.30. Negative β€” the short side fills in the long side's losing months, an internal hedge. Combined volatility 15.1% vs 17.8% if uncorrelated.

MAE / MFE Analysis (stop-loss & take-profit placement)

MetricValueDescription
Avg MAE β€” winners0.64%avg adverse move after entry
Avg MAE β€” losers1.52%2.4Γ— that of winners
Separation (Cohen's d)1.130.8+ = clear separation β†’ stop line is valid
Suggested stop (winners 90th pct)1.49%cuts 47% of losers earlier
MFE capture (winners)69%share of peak unrealized profit actually banked
Holding–return corr ρ0.63positive = trend-following behavior
Winners move only 0.64% against entry on average, while losers go 1.52% (2.4Γ— the difference). Placing the stop near 1.5% would preserve 90% of winners while cutting 47% of losers earlier than now.

πŸ’° Profit Share β€” 20%, Paid Per Trade

Total payout (fixed principal)
38.4K USDT
399 payouts
Total payout (compounded)
18.76M USDT
when payouts leave the principal
Effective rate vs net profit
30.2%
vs 20% nominal
Largest single payout
598 USDT
My net profit after payout
88.7K USDT
before payout 127.1K

Equity & Cumulative Payout (month-end, log scale)

10^510^610^710^8 2019202120222023202420252026 Equity USDT Cum. payout
Equity before payoutEquity after payoutCumulative payout

Metrics β€” Before vs After Payout (compounded equity)

ScenarioCumulative returnAnnualizedSharpeProfit factorMDDFinal equity
β‘  Before payout8,950,853.6%450.8%2.921.89-24.61%895.10M USDT
β‘‘ After per-trade payout296,059.7%230.7%2.491.65-27.65%29.62M USDT

Payouts by Year (fixed principal, USDT)

827,7717,0926,0283,0545,9316,2562,14920192020202120222023202420252026
YearNet profit (pre-payout)Per-trade payoutKept
2019273 USDT82 USDT191 USDT
202029.3K USDT7,771 USDT21.5K USDT
202125.8K USDT7,092 USDT18.7K USDT
202221.9K USDT6,028 USDT15.9K USDT
20235,761 USDT3,054 USDT2,707 USDT
202419.9K USDT5,931 USDT14.0K USDT
202520.7K USDT6,256 USDT14.5K USDT
20263,444 USDT2,149 USDT1,296 USDT
The nominal rate is 20%, but the effective rate on net profit is 30.2%. Per-trade settlement does not net losses β€” you share the wins and eat the losses alone β€” so the effective burden exceeds the nominal rate. Under compounding, payouts also take away principal that would have kept compounding, so the final-equity gap widens far beyond the rate (895.10M β†’ 29.62M USDT).

πŸ“‹ Validation Checklist (21 items)

βœ… Passed
20
⚠️ Warning
0
🚨 Caution
0
ℹ️ Info
1
Items validated
21
computable items only
πŸ“ 1. Data Quality
1. Data Integrity Check#1βœ… Passed
Exhaustively checks data-quality issues such as missing values (NaN), infinities (Inf), duplicate rows, and out-of-order dates. Many missing values distort every metric β€” Sharpe, MDD, and more β€” making the entire validation untrustworthy.
πŸ“ Benchmark 0 missing β†’ perfect βœ… | 1–5 β†’ inspect and fix ⚠️ | 5+ β†’ strategy needs re-review 🚨
Result 0 missing / 1846 rows Γ— 17 cols
πŸ’‘ Interpretation 0 missing values β€” the data is complete, so every statistic below was computed from the raw records. Results are trustworthy without preprocessing distortion.
πŸ“ 2. Model & Parameter Risk
2. Walk-Forward Analysis#3βœ… Passed
Splits the full period into rolling (fixed-window) and anchored (expanding-window) segments and repeatedly checks whether performance in each in-sample (IS) segment carries into the next out-of-sample (OOS) segment. WFE (Walk-Forward Efficiency) = mean OOS Sharpe ÷ mean IS Sharpe. Embargo buffer trades between segments prevent adjacent-regime information leakage. ⚠️ Trade-history CSVs contain no parameters, so this version measures performance consistency only, without re-optimization.
πŸ“ Benchmark WFE β‰₯ 60% β†’ robust βœ… | 40–60% β†’ moderate ⚠️ | below 40% or many OOS losses β†’ suspect overfitting / strategy decay 🚨
Result WFE rolling 83% / anchored 74% | profitable OOS 6/6 | embargo 18 trades
πŸ’‘ Interpretation Rolling through 6 steps, 6/6 OOS segments were profitable, with WFE of 83% rolling / 74% anchored. Above 60% means the training-period edge largely reappears in the next segment β€” good temporal robustness. An embargo of 18 trades between train and test blocks adjacent-regime leakage.
3. Out-of-Sample Test (70/30)#4βœ… Passed
Final validation on future data never used in strategy development. Going live on in-sample performance alone, without OOS validation, is a common recipe for disappointment. This tool automatically splits the trade history chronologically 70% (IS) / 30% (OOS) β€” see the πŸ§ͺ Deep Validation tab for detailed curves.
πŸ“ Benchmark OOS return β‰₯ 60% of IS β†’ good βœ… | 30–60% β†’ caution ⚠️ | below 30% or negative β†’ unfit for live trading 🚨
Result IS Sharpe 3.21 β†’ OOS 2.27 (retention 71%) | IS 936,699.5% / OOS 855.5%
πŸ’‘ Interpretation Sharpe in the untouched last 30% is 2.27, i.e. 71% of the first 70% (3.21). Retention above 60% means the edge survives market change β€” overfitting is unlikely.
4. PBO (Probability of Backtest Overfitting)#33βœ… Passed
CSCV (Combinatorially Symmetric Cross-Validation): splits the data into 6 blocks, forms C(6,3)=20 IS/OOS combinations, and measures how often the in-sample winner loses out-of-sample. 50% is coin-flip level β€” pure luck. (Accuracy degrades with few trades.)
πŸ“ Benchmark PBO < 10% β†’ no overfitting βœ… | 10–25% β†’ low risk βœ… | 25–50% β†’ caution ⚠️ | over 50% β†’ overfitting confirmed 🚨
Result PBO 0.0% | 0 of 20 splits lost OOS | purge/embargo 18 trades/boundary
πŸ’‘ Interpretation In 0% of the 20 train/test combinations, the in-sample winner lost out-of-sample. Low β€” a structural edge that holds however the period is sliced.
πŸ“ 3. Stress Tests
5. Monte Carlo Path Risk#5βœ… Passed
Simulates thousands of equity paths by randomly reshuffling trade order. If the actual result sits near the top of the distribution, luck likely played a role. The 95% worst-case MDD shows the realistic loss boundary.
πŸ“ Benchmark 95% worst MDD within βˆ’25% β†’ safe βœ… | βˆ’25% to βˆ’40% β†’ caution ⚠️ | beyond βˆ’40% β†’ high risk 🚨
Result 95% worst MDD -31.97% | 95% final return [865,429%, 107,127,486%]
πŸ’‘ Interpretation Reshuffling the same trades 1000 times gives a 95% worst MDD of -32.0%, which is deeper than the -24.6% actually experienced β€” the real path was a lucky ordering. Set margin and psychological stop limits to the 95% value (-32.0%), not the realized MDD.
6. Historical Stress Test#6βœ… Passed
Tests resilience during extreme drawdown periods such as the 2020 COVID crash (βˆ’50%) and the 2022 LUNA collapse. Strategies that look fine in normal times often break in crises. This tool auto-validates against the actual date ranges of COVID, the May-2021 crash, LUNA, FTX, the 2024 yen-carry unwind, the 2025 tariff shock, and the Oct-2025 mass liquidation β€” details in the πŸ§ͺ Deep Validation tab.
πŸ“ Benchmark Crisis MDD ≀ 2Γ— normal-period MDD β†’ robust βœ… | 2–3Γ— β†’ caution ⚠️ | over 3Γ— β†’ crisis-fragile 🚨
Result COVID crash: +70.1% | May-2021 crash: +19.7% | LUNA collapse: +33.9% | FTX bankruptcy: +23.7% | Yen-carry unwind: +34.7% | Tariff shock: +50.5% | Oct-2025 mass liquidation: +17.4%
πŸ’‘ Interpretation Judged on 'exposed trades' whose holding period overlapped a crisis. Of 7 exposed windows, 7 were defended successfully.
7. Transaction Cost Impact#7βœ… Passed
Checks whether profits survive after deducting commissions, slippage, and spread. Scalping and high-frequency strategies are extremely cost-sensitive β€” a 0.1% fee difference can flip annual returns.
πŸ“ Benchmark Profit after costs > 0 β†’ pass βœ… | profit shrinks 50%+ β†’ consider revising ⚠️ | turns negative β†’ unfit for live trading 🚨
Result commission 0.04%/side + slippage 0.00% (923 trades Γ— 1Γ—) | before 18,654,023.4% β†’ after 8,950,853.6%
πŸ’‘ Interpretation At 0.04% per side (round-trip 0.08%) Γ— 1Γ— leverage, the cumulative return is 8,950,853.6% β€” still profitable: a live-worthy edge that beats cost friction. Costs consumed 9,703,169.8pp.
πŸ“ 4. Performance Metrics
8. Kelly Criterion Allocation#9βœ… Passed
Kelly formula: f* = win rate βˆ’ (loss rate Γ· payoff ratio). The theoretically optimal bet fraction that maximizes long-run geometric growth. Full Kelly is extremely volatile, so half Kelly (50%) or quarter Kelly (25%) is the practical standard.
πŸ“ Benchmark Kelly > 0 β†’ positive expectancy βœ… | half Kelly 5–15% β†’ stable βœ… | 15–30% β†’ aggressive ⚠️ | Kelly ≀ 0 β†’ losing strategy 🚨
Result Full Kelly 28.64% | Half Kelly 14.32%
πŸ’‘ Interpretation Risking 28.6% of capital per trade mathematically maximizes long-run compound growth. Full Kelly's drawdowns are extreme, though β€” start live at half Kelly (14.3%) and adjust.
9. Win Rate & Win/Loss Ratio#10ℹ️ Info
Win rate and payoff ratio complement each other: a low win rate can still profit with a high payoff (trend-following), and a low payoff can profit with a high win rate (market-making). What matters is expectancy = (win rate Γ— avg win) βˆ’ (loss rate Γ— avg loss) > 0.
πŸ“ Benchmark Expectancy > 0 required 🚨 | WR β‰₯ 50% + payoff β‰₯ 1.3 β†’ good βœ… | WR 40–50% + payoff β‰₯ 2.0 β†’ good trend-follower βœ… | expectancy ≀ 0 β†’ not tradable
Result Win rate 43.2% | payoff 3.89 | W 399 / L 524 / total 923
πŸ’‘ Interpretation Win rate 43.2% Γ— payoff 3.89 β†’ expectancy +1.38% per trade. Classic trend-following profile β€” loses often but wins big. Losing streaks are inevitable, so money management to survive them is key.
10. Sharpe Ratio#11βœ… Passed
Sharpe = (mean return βˆ’ risk-free rate) Γ· return std dev Γ— √(trades per year). The standard performance yardstick for hedge funds and institutions. High volatility can drag it down even with high returns; crypto's volatility means different thresholds than equities.
πŸ“ Benchmark < 1.0 β†’ unfit 🚨 | 1.0–1.5 β†’ moderate ⚠️ | 1.5–2.0 β†’ good βœ… | 2.0–3.0 β†’ excellent βœ… | above 3.0 β†’ outstanding (re-check for overfitting)
Result Annualized Sharpe 2.92 (1.5+ good / 2.0+ excellent)
πŸ’‘ Interpretation The strategy earned 2.92 units of excess return per unit of risk (volatility). Excellent β€” meets institutional standards.
11. Sortino Ratio#12βœ… Passed
Sortino = (mean return βˆ’ risk-free rate) Γ· downside std dev Γ— √(trades per year). Unlike Sharpe, upside volatility isn't treated as risk. Better suited to trend-following and long-biased strategies; Sortino > Sharpe indicates a favorably asymmetric profit structure.
πŸ“ Benchmark < 1.5 β†’ unfit 🚨 | 1.5–2.0 β†’ moderate ⚠️ | 2.0–3.0 β†’ good βœ… | above 3.0 β†’ excellent βœ… | Sortino > Sharpe β†’ positive asymmetry βœ…
Result Annualized Sortino 15.97
πŸ’‘ Interpretation Sortino 15.97 clearly exceeds Sharpe (2.92) β€” most volatility comes from the profit side. Losses are small and uniform while gains spike: a desirable asymmetry for risk management.
12. Calmar Ratio#13βœ… Passed
Calmar = annualized return Γ· |max drawdown|. Shows how much you earn per year relative to the worst loss period you had to sit through. A key screening criterion for long-horizon CTA funds; a larger MDD lowers Calmar, cleanly reflecting return-per-risk efficiency.
πŸ“ Benchmark < 0.5 β†’ unfit 🚨 | 0.5–1.0 β†’ marginal ⚠️ | 1.0–2.0 β†’ moderate | 2.0–3.0 β†’ good βœ… | above 3.0 β†’ excellent βœ…
Result Calmar 18.31 (annual return 450.8% / MDD -24.61%)
πŸ’‘ Interpretation Earning 450.8%/yr against a worst drawdown of -24.6% β€” one year's return covers the MDD about 18.3 times. The worst drawdown is quickly repaid by a year's returns.
13. Profit Factor#14βœ… Passed
Profit Factor = gross profit Γ· gross loss. The most intuitive profitability metric: 1.0 is break-even, 1.5 means $1.50 earned per $1 lost. Best checked both before and after trading costs.
πŸ“ Benchmark < 1.0 β†’ losing strategy 🚨 | 1.0–1.3 β†’ marginal ⚠️ | 1.3–1.5 β†’ moderate | 1.5–2.0 β†’ good βœ… | above 2.0 β†’ excellent βœ…
Result Profit factor 2.05 (gross profit 3,631,551,095 / gross loss 1,769,562,261 USDT)
πŸ’‘ Interpretation Earned 2.05 USDT per 1 USDT lost. A strong structure earning 2Γ—+ per unit lost.
14. Recovery Factor#15βœ… Passed
Recovery Factor = net cumulative profit Γ· max drawdown amount. Shows how decisively the strategy overcomes its worst loss. Below 1.0 means the MDD hasn't been recovered yet; higher is better.
πŸ“ Benchmark < 1.0 β†’ MDD not recovered 🚨 | 1.0–2.0 β†’ marginal ⚠️ | 2.0–3.0 β†’ moderate | 3.0–5.0 β†’ good βœ… | above 5.0 β†’ excellent βœ…
Result Recovery factor 5.02 (net profit 1,861,988,833 / MDD 370,559,835 USDT)
πŸ’‘ Interpretation Net profit is 5.0Γ— the max-drawdown amount. Even repeated drawdowns leave overwhelming net profit.
15. Risk of Ruin#16βœ… Passed
Formula: Z = (loss rate Γ— loss size) Γ· (win rate Γ— avg win). Z < 1 means positive expectancy. RoR = Z^N, where N = ruin-threshold loss Γ· per-trade loss. The closer Z is to 1 and the smaller N is, the higher the ruin probability. Choose the loss-unit basis (average / max / custom) in the sidebar.
πŸ“ Benchmark RoR < 0.1% β†’ negligible βœ… | 0.1–1% β†’ low risk βœ… | 1–10% β†’ caution, review leverage ⚠️ | over 10% β†’ rethink position size now 🚨
Result Estimated ruin 0.0000% [avg loss 1.24%/trade / ruin threshold 50%]
πŸ’‘ Interpretation Under current conditions (win rate 43%, avg loss 1.24%/trade), losing 50% of capital requires roughly 40 losses' worth of drawdown, with probability 0.0000%. Effectively zero β€” the current position size is safe from a ruin standpoint.
16. Deflated Sharpe Ratio#32βœ… Passed
Bailey & LΓ³pez de Prado (2014): DSR = Ξ¦[(√(Nβˆ’1)Β·(SR βˆ’ E[SR*])) / √(1 βˆ’ γ₃·SR + (Ξ³β‚„βˆ’1)/4Β·SRΒ²)]. The probability that the observed Sharpe exceeds the expected maximum Sharpe (E[SR*]) implied by your parameter-search count. Enter the true number of backtest attempts (n_trials) in the sidebar for meaningful results.
πŸ“ Benchmark DSR β‰₯ 95% β†’ statistically significant βœ… | 80–95% β†’ moderate ⚠️ | below 80% β†’ suspect multiple-testing bias 🚨 | larger n_trials tightens the bar
Result DSR 100.0% | E[SR*] 0.0000 | attempts 1
πŸ’‘ Interpretation Attempts is set to 1, so no multiple-testing correction is applied (DSR 100%). If you backtested many parameter variants, entering the true count including discarded ones corrects strategy survivorship bias β€” reporting only the surviving winner.
πŸ“ 5. Market Regimes & External Factors
17. Market Regime Test#17βœ… Passed
Splits the backtest into thirds (early / middle / late) and compares win rate, mean return, and Sharpe across them. If only the early segment is good and performance decays later, the strategy is falling behind the market or its pattern has been arbitraged away.
πŸ“ Benchmark All 3 segments profitable + Sharpe spread ≀ 30% β†’ excellent βœ… | 1 losing segment β†’ caution ⚠️ | 2+ losing β†’ possible strategy decay 🚨
Result Early (1/3): 2.14% (WR 44%, SR 3.72) | Middle (2/3): 0.93% (WR 43%, SR 2.47) | Late (3/3): 1.07% (WR 43%, SR 2.53)
πŸ’‘ Interpretation All three segments profitable β€” the edge worked across the whole period rather than clustering in one era.
πŸ“ 6. Statistical Tests
18. Bootstrap Confidence Interval#20βœ… Passed
Bootstraps 1,000 resamples (with replacement) from the actual trades to compute a 95% confidence interval for the mean return. If the lower bound exceeds 0, profitability is statistically significant beyond sampling error. Fewer trades widen the interval.
πŸ“ Benchmark 95% CI lower bound > 0% β†’ statistically significant profit βœ… | lower bound ≀ 0% β†’ possibly luck ⚠️ | interval wider than the return itself β†’ insufficient sample 🚨
Result 95% CI [1.020%, 1.750%] (1,000 resamples)
πŸ’‘ Interpretation The 'true mean return' lies within [1.020%, 1.750%] with 95% confidence. The whole interval sits above 0% β€” statistically rejecting the possibility that the profits are a lucky sample arrangement.
19. Trade-Sequence Randomness (Runs Test)#23βœ… Passed
Runs test: checks whether win/loss streaks in the trade sequence are statistically more clustered than chance would produce. p > 0.05 means the sequence is consistent with randomness β€” no streak clustering. Pronounced streaks can make live results diverge sharply from the report.
πŸ“ Benchmark p > 0.10 β†’ clearly random βœ… | 0.05–0.10 β†’ borderline ⚠️ | p ≀ 0.05 β†’ streak pattern present, review structural bias 🚨
Result Actual runs 458 | expected 454.0 | Z=0.00 | p=0.9998
πŸ’‘ Interpretation p = 1.000 β€” the win/loss sequence is indistinguishable from random. No hidden streak clustering, so the Monte Carlo reshuffling results (#5) can be trusted as-is.
20. Time Delay Test#24βœ… Passed
Measures how Sharpe changes when entries are delayed by 1, 2, 3, or 5 trades after the signal. Delay sensitivity means the strategy depends on razor-thin entry timing that's hard to execute live; the bigger the slippage and order-book latency, the bigger the real-world degradation.
πŸ“ Benchmark Sharpe drop ≀ 10% after 1-trade delay β†’ robust βœ… | 10–30% β†’ delay-sensitive ⚠️ | over 30% or negative β†’ high live-execution risk 🚨
Result Original SR 2.92 | delay 1 β†’ 2.96 | delay 2 β†’ 2.96 | delay 3 β†’ 2.95 | delay 5 β†’ 2.95 | 1-trade-delay drop 1.2%
πŸ’‘ Interpretation Entering one trade late moves Sharpe 2.92 β†’ 2.96 (1% change). Performance holds even when fills lag a beat β€” a robust signal that tolerates alertβ†’API delays and retries.
21. Random Entry Test#25βœ… Passed
One-sided t-test: Hβ‚€ = mean return = 0. The lower the p-value, the more significantly the strategy beats random entry. A t-statistic above 3 is a strong signal, and more trades mean more confidence. Below 30 trades, statistical power drops sharply.
πŸ“ Benchmark T-stat β‰₯ 3.0 + p < 0.01 β†’ highly significant βœ… | T-stat β‰₯ 2.0 + p < 0.05 β†’ significant βœ… | p β‰₯ 0.05 β†’ not significant 🚨 | β‰₯ 100 trades recommended
Result T-stat 7.630 | one-sided p 0.0000
πŸ’‘ Interpretation T=7.63: the probability that a zero-mean strategy produces this performance by chance is 0.0000. Passes even the 1% significance level on 923 trades β€” this edge is very unlikely to be statistical noise.