Now in beta. Get free access before the public launch.

The Tests

Eight tests, each catching a different way a backtest can lie.

Every test comes from published research. We build the original algorithm, check it against the paper's own worked example, and run the full computation on your real trade log. Nothing is approximated to save time.

Scoring

What each test can cost you

Your audit starts at 100 points. Every test below can take points off, and the thresholds are fixed before we see your data. Here is the whole table in one place.

TestWhat it measuresFull penaltyPartial penalty
Probability of Backtest OverfittingPBO, 0 to 1−35 above 0.6−20 from 0.4 to 0.6
Walk-Forward DegradationDegradation ratio−35 below 0.2−25 below 0.5, −15 below 0.8
Minimum Sample SizeTrades vs minimum−30 below the minimumno partial credit
Monte Carlo Permutation Testp-value−25 above 0.10−10 from 0.05 to 0.10
Deflated Sharpe RatioPSR, 0 to 1−20 below 0.75−10 from 0.75 to 0.95
Regime AnalysisHHI, top-quarter share−20 at HHI 0.5 or 70% share−10 at HHI 0.3 or 50% share
Parameter SensitivityCoefficient of variation−15 above 0.5−7 from 0.25 to 0.5
Transaction Cost AnalysisNet vs gross returnno penaltyinformational only
Worked example. The retired strategy on the demo page lost 20 points on the overfitting check, 15 on walk-forward, and 20 on regime concentration. 100 − 20 − 15 − 20 = 45. That is a C. Your results page lists every deduction next to the number that caused it, so you can check the arithmetic yourself.
Test 01 / 08

Probability of Backtest Overfitting

The share of in-sample winners that turn into out-of-sample losers. A PBO above 50% means your best strategy was more likely picked by luck than by skill.

Did your strategy learn something real, or did it memorize the past? Bailey, Borwein, López de Prado and Zhu (2015) built a way to measure that directly. They call it Combinatorially Symmetric Cross-Validation, or CSCV. In plain terms: take the configurations you tested, split them into a training half and a testing half, and do it every way the halves can be drawn. For a 16-configuration upload that is C(N, N/2) = 12,870 splits.

In each split we find the configuration with the best in-sample Sharpe ratio, then look at where it ranks out-of-sample. PBO is the share of splits where the in-sample winner lands in the bottom half out-of-sample. A PBO of 0.0 means the winner always held up on new data. A PBO of 1.0 means it always fell over. Pure chance gives 0.5. Almost all retail backtests land between 0.5 and 0.9.

The scatter chart shows it directly. In-sample rank runs along the bottom, out-of-sample rank up the side. If the points hugged the diagonal, PBO would be 0. The highlighted point is the in-sample winner. Where it sits vertically tells you whether winning in training bought you anything.

Fourteen plain points are plotted with in-sample rank on the horizontal axis, 0 to 100, and out-of-sample rank on the vertical axis, 0 to 100. A dashed diagonal marks where the two ranks would agree perfectly. The points show no such pattern: at almost every in-sample rank there are points high and low on the out-of-sample scale. One highlighted point at the far right, labelled the in-sample winner and 8th out of sample, sits near the bottom of the plot. The annotation reads PBO equals 73 percent, high overfitting.Perfect correlationIS Rank (in-sample performance)OOS Rank01000100IS winnerOOS rank: 8thPBO = 73%High overfitting

Source paper

Bailey, D., Borwein, J., López de Prado, M. & Zhu, Q. (2015). The Probability of Backtest Overfitting. Journal of Computational Finance. SSRN 2326253 →

Pass example

PBO = 0.18
18% of in-sample winners lose out-of-sample. Good evidence the edge is real and not fitted.

Fail example

PBO = 0.73
73% of in-sample winners lose out-of-sample. The edge is most likely an artifact of picking a winner.

When this test can't tell you anything, we say so

One honest caveat. Full CSCV needs several parameter configurations in one upload. Upload a single strategy and we fall back to a bootstrap proxy. In our calibration on real trade logs, that proxy returned about 0.5 for every strategy, strong and weak alike (0.467 to 0.501 across all six). So in single-strategy mode it cannot tell you much, and we flag it as a limited-information result instead of pretending otherwise. No other tool tells you when its own test is useless. Multi-configuration uploads run the real CSCV computation and are not affected. See the calibration report →

Score impact: Up to −35 points (highest weight, tied with Walk-Forward). PBO above 0.6 takes the full −35. PBO between 0.4 and 0.6 takes −20.
Test 02 / 08

Walk-Forward Degradation

Out-of-sample Sharpe divided by in-sample Sharpe. It measures how far performance falls when the strategy meets data it has never seen.

We hold out the final 30% of your backtest. Whatever you tuned during development, you did not tune it on that stretch. The first 70% is the in-sample period, the data you worked with. The degradation ratio is out-of-sample Sharpe divided by in-sample Sharpe.

A ratio above 0.8 means the strategy keeps at least 80% of its performance on new data. That is acceptable and takes no penalty. Between 0.5 and 0.8 is a moderate warning. Between 0.2 and 0.5 the drop is significant. Below 0.2 is critical: most of the performance was never real, which is the signature of in-sample overfitting.

In the chart, the in-sample curve is indigo and climbing. The out-of-sample curve past the divider is red and falling. Most retail backtests lose 70–85% of their Sharpe the first time they see a walk-forward test. The edge evaporates the moment the data is unfamiliar.

Walk-Forward Analysis · EMA_Cross_v3
In-sample (70%)Out-of-sample (30%)
A dashed vertical divider separates the left 70 percent of the chart, labelled in-sample, from the right 30 percent, labelled out-of-sample. The solid line climbs without interruption from 0 percent to roughly 70 percent across the in-sample period. From the divider onward the dashed line turns down and falls back to under 20 percent by the end of the out-of-sample period. The annotation records Sharpe falling from 2.1 to 0.4, a degradation of 81 percent.0%20%40%60%IN-SAMPLEOUT-OF-SAMPLESharpe 2.10.481% degradation

Foundation

Pardo, R. (2008). The Evaluation and Optimization of Trading Strategies. Wiley. Walk-forward validation framework with rolling windows.

Pass example

Ratio: 0.84
Out-of-sample Sharpe is 84% of in-sample. An acceptable drop. The edge looks stable.

Fail example

Ratio: 0.19
Out-of-sample Sharpe is 19% of in-sample. 81% of the performance you were counting on is gone.
Score impact: Up to −35 points (highest weight, tied with PBO). A degradation ratio below 0.2 takes the full −35; between 0.2 and 0.5 takes −25; between 0.5 and 0.8 takes −15.
Test 03 / 08

Monte Carlo Permutation Test

How often a randomly shuffled version of your trades beats the real thing. A high p-value means luck could have produced your returns.

Could pure luck have produced your result? The permutation test asks exactly that. We shuffle the order of your trades 10,000 times and count how often a random ordering beats you.

Each shuffle gets its own Sharpe ratio. The share of shuffles that beat your actual Sharpe is your p-value. A p-value of 0.04 means only 4% of random orderings beat you, which is significant at the 5% level. A p-value of 0.31 means 31% beat you. At that point, random ordering explains your returns about as well as skill does.

This test is deliberately blind to the cherry-picking that PBO catches. It asks whether your individual trades show skill, whatever parameter set produced them. A strategy can pass PBO and fail Monte Carlo, or the other way around. That is a feature, not a contradiction.

Monte Carlo · 10,000 random trade orderings vs. actual strategy
Sixteen bars run from low Sharpe on the left to high Sharpe on the right, rising to a peak just past the middle and tapering away at both ends. A dashed vertical line marks where the actual strategy falls, on the low side of that peak and well short of the right tail. The annotation states that 31 percent of the random orderings beat the strategy, for a p-value of 0.31, labelled Not significant.Your strategy31% beat your strategyLow SharpeHigh SharpeFreqp-value: 0.31Not significant

Foundation

Efron, B. & Tibshirani, R. (1993). An Introduction to the Bootstrap. Chapman & Hall. Permutation testing framework applied to return sequences.

Pass example

p-value: 0.03
Only 3% of random orderings beat the strategy. Statistically significant.

Fail example

p-value: 0.31
31% of random orderings beat the strategy. Random ordering explains the returns.
Score impact: Up to −25 points. A p-value above 0.10 takes the full −25; between 0.05 and 0.10 takes −10.
Test 04 / 08

Deflated Sharpe Ratio

Your Sharpe ratio, corrected for how many configurations you tested. More tries raise the bar a result has to clear to mean anything.

Harvey, Liu and Zhu (2016) found that a Sharpe ratio needs to be above 2.0 before it should survive academic review, because researchers typically test hundreds of configurations before they submit anything. Bailey and López de Prado (2014) derived the correction. The Deflated Sharpe Ratio adjusts for the best Sharpe you would expect from random configurations, given how many you tried.

The chart shows the Sharpe you need against the number of configurations tested, on a log scale. Test 25 configurations and report a Sharpe of 1.4, and you are below the line. One of those 25 could plausibly have produced that number by chance alone.

The test uses the trial count you enter on the upload form. Report 1 configuration, meaning you built the strategy once and ran it, and any Sharpe above 1.0 passes. Report 100 and you need a Sharpe above 2.3. That is the single best reason to answer the trials field honestly. Understating it makes this test lenient, and the leniency is not doing you any favors.

The horizontal axis counts configurations tested on a log scale, marked 1, 5, 10, 25 and 50; the vertical axis is the required Sharpe, marked 0.5 through 2.0. The rising curve starts near 1.0 at a single configuration and climbs past 2.0 by 50 configurations. A dashed horizontal line marks a strategy Sharpe of 1.4. The two cross early on, and the whole area to the right of that crossing and below the strategy line is shaded and labelled as failing the threshold above about six configurations.0.51.01.52.015102550Configurations tested (log scale)Required SharpeYour Sharpe: 1.4Fails threshold above ~6 configsRequired SRthreshold

Source paper

Bailey, D. & López de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality. Journal of Portfolio Management. SSRN 2460551 →

Pass example

DSR: 1.14
Reported Sharpe 1.8 clears the threshold of 1.58 for 10 configurations tested. Passes.

Fail example

DSR: 0.61
Reported Sharpe 1.4 sits under the threshold of 2.30 for 50 configurations tested. Fails.
Score impact: Up to −20 points. A Probabilistic Sharpe Ratio below 0.75 takes the full −20; between 0.75 and 0.95 takes −10.
Test 05 / 08

Regime Analysis

Your returns split by market regime. A strategy that only profits in one regime is a bet on that condition continuing, not evidence of an edge.

Markets alternate between three conditions, and each one rewards different behavior. Trending markets run in one direction and pay momentum. Ranging markets chop around a mean and pay mean reversion. High-volatility markets punish nearly everyone. We label each quarter of your backtest with a Hidden Markov Model fitted to volatility and autocorrelation.

A real edge should make money in at least two of the three. A strategy that only profits in trending quarters looks like it works across market history, but only because most of the profitable stretches happened to be trending. When the regime flips, so does the strategy.

The bar chart shows quarterly P&L colored by regime. Steady indigo bars with the occasional amber loss are fine. Consistent red losses in high-volatility quarters mean the strategy breaks exactly when market risk peaks. That is the worst version of this problem.

Bars run above and below a zero line on a scale from plus 10 to minus 10 percent. The five trending-regime bars are all positive, between about 2 and 12 percent. The four ranging-regime bars are small, three of them losses of 1 to 4 percent and one a gain of about 3 percent. The three high-volatility bars are all losses, at roughly minus 6, minus 9 and minus 7 percent, and the annotation reads: losing in high-vol regimes.+10%0%−10%TrendingRangingHigh-volLosing inhigh-vol regimes

Methodology

Hidden Markov Model regime labeling following Ang & Timmermann (2012). Regime Changes and Financial Markets. Annual Review of Financial Economics.

Pass example

3/3 regimes
Positive returns in trending, ranging, and high-vol regimes. The edge does not need one market mood.

Fail example

1/3 regimes
Profitable in trending markets only. Loses in ranging and high-vol. A bet on one condition.
Score impact: Up to −20 points. An HHI above 0.5, or more than 70% of profits concentrated in the top quarter of periods, takes the full −20. Moderate concentration (HHI above 0.3, or top-quarter share above 50%) takes −10.
Test 06 / 08

Minimum Sample Size

The minimum number of trades your claims need behind them, given how many configurations you tested.

Bailey and López de Prado (2014) worked out the Minimum Backtest Length: how many observations a backtest needs before its Sharpe ratio means anything, given the Sharpe you report and the number of configurations you tried. More tries, more trades needed for the same confidence.

The chart shows three curves: the minimum Sharpe required against trade count, for 1, 10, and 100 configurations tested. Where you sit on it decides whether you have enough evidence. A strategy with 47 trades and a Sharpe of 1.4, tested across 10 configurations, sits just above the line. It passes, barely. The same Sharpe on 20 trades would not.

The practical version: 30 trades is not enough for most real strategies. With 30 trades and 50 configurations tested, you need a Sharpe above 2.1 to pass. Very few retail backtests clear that bar honestly.

The horizontal axis counts trades from 10 to 200 and the vertical axis the minimum required Sharpe, from 0 to above 2. Three curves are drawn, one solid for a single configuration and two dashed for 10 and 100 configurations. Each begins high at 10 trades and drops sharply before levelling out past 100 trades, and the 100-configuration curve stays above the other two across the whole range. A highlighted point at 47 trades and a Sharpe of about 1.4 sits above all three curves and is labelled as passing.01.02.01050100200Trade countMin required SharpeYour strategypasses (47 trades)1 config10 configs100 configs

Source paper

Bailey, D. & López de Prado, M. (2014). The Deflated Sharpe Ratio. Minimum Backtest Length derivation. SSRN 2460551 →

Pass example

142 trades
Well above the minimum for the reported Sharpe and trial count.

Fail example

22 trades
Below the minimum. The other tests are underpowered here. Do not lean on the results.
Score impact: −30 points on failure. Falling below the minimum required trade count for your reported Sharpe and trial count takes the full penalty. There is no partial credit for underpowered evidence.
Test 07 / 08

Transaction Cost Analysis

Your returns recomputed with realistic slippage and commission. Costs look small per trade and pile up across hundreds of them into a serious drag.

A backtest that fills market orders at the close or the mid price is optimistic. Real fills pay slippage, which is the gap between the price you expected and the price you got. They also pay commission and the bid-ask spread, on the way in and the way out. A strategy trading 200 times a year at 0.2% slippage per trade picks up 40% of extra drag annually.

We apply cost estimates by asset class. Futures: 0.05–0.1% slippage plus $4.50 round-trip commission per contract. Forex: 1.5–2 pip spread equivalent. Crypto: 0.1% taker fee plus 0.1–0.3% slippage on liquid pairs, more on thin ones. Equities: 0.02–0.05% slippage plus $0.005 per share commission.

The two equity curves show what execution takes. A strategy with a 52% gross return and a 12% net return is not a strategy. The apparent edge was the cost you were not paying.

The vertical axis is account value, marked 10, 12 and 14 thousand dollars. A solid line traces equity before costs, rising steeply through the middle of the period and flattening near the top of the chart at plus 52 percent. A dashed line traces the same strategy after realistic trading costs; it follows the same shape but much shallower, ending at plus 12 percent. The two start together and the gap between them widens across the whole period.$10k$12k$14k+52% (no costs)+12% (with costs)Gross returnsAfter realistic costs

Methodology

Cost estimates based on industry-standard slippage models following Almgren et al. (2005). Direct Estimation of Equity Market Impact. Risk Magazine.

Pass example

Net: +38%
Gross return 44%, net after costs 38%. Costs are manageable. The strategy survives them.

Fail example

Net: −4%
Gross return 22%, net after realistic costs −4%. The strategy only works if trading is free.
Score impact: None. This test is informational. Broker costs vary too widely for a fair universal penalty, so we show the cost-adjusted curve and attach a high-cost warning to your recommendation when realistic costs cut your Sharpe ratio by more than 40%. Your score is unaffected. Your decision should not be.
Test 08 / 08

Parameter Sensitivity

How far performance moves when parameters shift by ±20%. A real edge is stable across a range. A narrow spike is a curve fitted to history.

A real trading edge exists because of something structural in the market, not because your RSI period is exactly 14 rather than 13 or 15. If your strategy needs exact numbers to perform, those numbers were fitted to the data you had. They are not evidence of anything repeatable.

We move each parameter by ±20% of the value you reported and compute the Sharpe ratio at each point. The sensitivity score is the coefficient of variation of those Sharpes, which is a plain measure of how much they bounce around. A flat plateau scores well. A sharp peak with steep sides triggers a warning.

The chart shows Sharpe against RSI period for a strategy fitted to RSI 14. A sharp peak at 14, with Sharpe dropping 78% at 12 or 16, is the classic overfitting shape. Strategies that look like this almost never survive live trading.

The RSI period parameter runs along the horizontal axis from 5 to 30 and the Sharpe ratio up the vertical axis from 0 to above 2. The curve climbs from about 0.5 at period 5 to a sharp peak of 1.62 at period 14, the value marked by a dashed vertical line as the one actually tested, then falls away steadily to roughly 0.2 by period 30. The annotation notes that the Sharpe drops 78 percent outside the band from RSI 12 to 16.01.02.05152530RSI period parameterRSI=14 testedSR: 1.62SR drops 78%outside RSI 12–16

Methodology

Sensitivity analysis following Ioannidis (2005) principles for stability checking. Parameter space exploration using ±20% variation with Sharpe ratio as the objective function.

Pass example

CV: 0.08
Sharpe moves less than 8% across the ±20% range. The edge does not depend on exact values.

Fail example

CV: 0.74
Sharpe moves 74% across the ±20% range. Performance hangs on the exact numbers.
Score impact: Up to −15 points. A coefficient of variation above 0.5 takes the full −15; between 0.25 and 0.5 takes −7.
Validation

Calibrated against ground truth

In August 2026 we ran these eight tests on six real strategies whose quality was already known: a live, validated edge, its retired and degraded variants, and two strategies the research program that built them had killed. The engine was not told which was which.

calibration_run.txt
// six real MetaTrader 5 backtests of known quality, Aug 2026
// labels below are the ground truth, not an engine output
 
strategy                  trades      PF     score   runtime
live incumbent               867    1.30    60 / B     20.7s
clock-corrected twin         898    1.28    60 / B     23.2s
degraded variant           1,111    1.18    60 / B     20.7s
retired variant              581    1.24    45 / C     16.8s
killed fade                  284    1.11     0 / F     13.4s
killed strategy              164    0.62     0 / F     12.6s
 
// the two killed strategies graded F. the live edge graded B.
// slowest audit 23.2s, fastest 12.6s

Result

The killed strategies graded F. The live edge graded B and beat everything that had been killed. Sorting all six by the engine's continuous outputs, the Monte Carlo p-value and the Probabilistic Sharpe Ratio, reproduces the known quality ordering exactly. Every audit finished in under 30 seconds. The report also lists what the engine gets wrong, because a methodology page that hides its limits is not a methodology page. Read the full calibration report →

What this run does not prove. The score is a verdict on your evidence, not a forecast. Two of these six sit at the same 60 / B while the underlying tests separate them correctly, because the penalty bands are coarse. We list that limit, and three others, in the report rather than in a footnote nobody reads.

Most strategies fail these tests. Yours probably will too.

Better to know now, while it's free, than after you've funded an account.

Start free, no credit card