Every test comes from published research. We build the original algorithm, check it against the paper's own worked example, and run the full computation on your real trade log. Nothing is approximated to save time.
Your audit starts at 100 points. Every test below can take points off, and the thresholds are fixed before we see your data. Here is the whole table in one place.
01 Overfitting · 02 Walk-forward · 03 Monte Carlo · 04 Deflated Sharpe · 05 Regime · 06 Sample size · 07 Costs · 08 Sensitivity · Calibration
The share of in-sample winners that turn into out-of-sample losers. A PBO above 50% means your best strategy was more likely picked by luck than by skill.
Did your strategy learn something real, or did it memorize the past? Bailey, Borwein, López de Prado and Zhu (2015) built a way to measure that directly. They call it Combinatorially Symmetric Cross-Validation, or CSCV. In plain terms: take the configurations you tested, split them into a training half and a testing half, and do it every way the halves can be drawn. For a 16-configuration upload that is C(N, N/2) = 12,870 splits.
In each split we find the configuration with the best in-sample Sharpe ratio, then look at where it ranks out-of-sample. PBO is the share of splits where the in-sample winner lands in the bottom half out-of-sample. A PBO of 0.0 means the winner always held up on new data. A PBO of 1.0 means it always fell over. Pure chance gives 0.5. Almost all retail backtests land between 0.5 and 0.9.
The scatter chart shows it directly. In-sample rank runs along the bottom, out-of-sample rank up the side. If the points hugged the diagonal, PBO would be 0. The highlighted point is the in-sample winner. Where it sits vertically tells you whether winning in training bought you anything.
Bailey, D., Borwein, J., López de Prado, M. & Zhu, Q. (2015). The Probability of Backtest Overfitting. Journal of Computational Finance. SSRN 2326253 →
One honest caveat. Full CSCV needs several parameter configurations in one upload. Upload a single strategy and we fall back to a bootstrap proxy. In our calibration on real trade logs, that proxy returned about 0.5 for every strategy, strong and weak alike (0.467 to 0.501 across all six). So in single-strategy mode it cannot tell you much, and we flag it as a limited-information result instead of pretending otherwise. No other tool tells you when its own test is useless. Multi-configuration uploads run the real CSCV computation and are not affected. See the calibration report →
Out-of-sample Sharpe divided by in-sample Sharpe. It measures how far performance falls when the strategy meets data it has never seen.
We hold out the final 30% of your backtest. Whatever you tuned during development, you did not tune it on that stretch. The first 70% is the in-sample period, the data you worked with. The degradation ratio is out-of-sample Sharpe divided by in-sample Sharpe.
A ratio above 0.8 means the strategy keeps at least 80% of its performance on new data. That is acceptable and takes no penalty. Between 0.5 and 0.8 is a moderate warning. Between 0.2 and 0.5 the drop is significant. Below 0.2 is critical: most of the performance was never real, which is the signature of in-sample overfitting.
In the chart, the in-sample curve is indigo and climbing. The out-of-sample curve past the divider is red and falling. Most retail backtests lose 70–85% of their Sharpe the first time they see a walk-forward test. The edge evaporates the moment the data is unfamiliar.
Pardo, R. (2008). The Evaluation and Optimization of Trading Strategies. Wiley. Walk-forward validation framework with rolling windows.
How often a randomly shuffled version of your trades beats the real thing. A high p-value means luck could have produced your returns.
Could pure luck have produced your result? The permutation test asks exactly that. We shuffle the order of your trades 10,000 times and count how often a random ordering beats you.
Each shuffle gets its own Sharpe ratio. The share of shuffles that beat your actual Sharpe is your p-value. A p-value of 0.04 means only 4% of random orderings beat you, which is significant at the 5% level. A p-value of 0.31 means 31% beat you. At that point, random ordering explains your returns about as well as skill does.
This test is deliberately blind to the cherry-picking that PBO catches. It asks whether your individual trades show skill, whatever parameter set produced them. A strategy can pass PBO and fail Monte Carlo, or the other way around. That is a feature, not a contradiction.
Efron, B. & Tibshirani, R. (1993). An Introduction to the Bootstrap. Chapman & Hall. Permutation testing framework applied to return sequences.
Your Sharpe ratio, corrected for how many configurations you tested. More tries raise the bar a result has to clear to mean anything.
Harvey, Liu and Zhu (2016) found that a Sharpe ratio needs to be above 2.0 before it should survive academic review, because researchers typically test hundreds of configurations before they submit anything. Bailey and López de Prado (2014) derived the correction. The Deflated Sharpe Ratio adjusts for the best Sharpe you would expect from random configurations, given how many you tried.
The chart shows the Sharpe you need against the number of configurations tested, on a log scale. Test 25 configurations and report a Sharpe of 1.4, and you are below the line. One of those 25 could plausibly have produced that number by chance alone.
The test uses the trial count you enter on the upload form. Report 1 configuration, meaning you built the strategy once and ran it, and any Sharpe above 1.0 passes. Report 100 and you need a Sharpe above 2.3. That is the single best reason to answer the trials field honestly. Understating it makes this test lenient, and the leniency is not doing you any favors.
Bailey, D. & López de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality. Journal of Portfolio Management. SSRN 2460551 →
Your returns split by market regime. A strategy that only profits in one regime is a bet on that condition continuing, not evidence of an edge.
Markets alternate between three conditions, and each one rewards different behavior. Trending markets run in one direction and pay momentum. Ranging markets chop around a mean and pay mean reversion. High-volatility markets punish nearly everyone. We label each quarter of your backtest with a Hidden Markov Model fitted to volatility and autocorrelation.
A real edge should make money in at least two of the three. A strategy that only profits in trending quarters looks like it works across market history, but only because most of the profitable stretches happened to be trending. When the regime flips, so does the strategy.
The bar chart shows quarterly P&L colored by regime. Steady indigo bars with the occasional amber loss are fine. Consistent red losses in high-volatility quarters mean the strategy breaks exactly when market risk peaks. That is the worst version of this problem.
Hidden Markov Model regime labeling following Ang & Timmermann (2012). Regime Changes and Financial Markets. Annual Review of Financial Economics.
The minimum number of trades your claims need behind them, given how many configurations you tested.
Bailey and López de Prado (2014) worked out the Minimum Backtest Length: how many observations a backtest needs before its Sharpe ratio means anything, given the Sharpe you report and the number of configurations you tried. More tries, more trades needed for the same confidence.
The chart shows three curves: the minimum Sharpe required against trade count, for 1, 10, and 100 configurations tested. Where you sit on it decides whether you have enough evidence. A strategy with 47 trades and a Sharpe of 1.4, tested across 10 configurations, sits just above the line. It passes, barely. The same Sharpe on 20 trades would not.
The practical version: 30 trades is not enough for most real strategies. With 30 trades and 50 configurations tested, you need a Sharpe above 2.1 to pass. Very few retail backtests clear that bar honestly.
Bailey, D. & López de Prado, M. (2014). The Deflated Sharpe Ratio. Minimum Backtest Length derivation. SSRN 2460551 →
Your returns recomputed with realistic slippage and commission. Costs look small per trade and pile up across hundreds of them into a serious drag.
A backtest that fills market orders at the close or the mid price is optimistic. Real fills pay slippage, which is the gap between the price you expected and the price you got. They also pay commission and the bid-ask spread, on the way in and the way out. A strategy trading 200 times a year at 0.2% slippage per trade picks up 40% of extra drag annually.
We apply cost estimates by asset class. Futures: 0.05–0.1% slippage plus $4.50 round-trip commission per contract. Forex: 1.5–2 pip spread equivalent. Crypto: 0.1% taker fee plus 0.1–0.3% slippage on liquid pairs, more on thin ones. Equities: 0.02–0.05% slippage plus $0.005 per share commission.
The two equity curves show what execution takes. A strategy with a 52% gross return and a 12% net return is not a strategy. The apparent edge was the cost you were not paying.
Cost estimates based on industry-standard slippage models following Almgren et al. (2005). Direct Estimation of Equity Market Impact. Risk Magazine.
How far performance moves when parameters shift by ±20%. A real edge is stable across a range. A narrow spike is a curve fitted to history.
A real trading edge exists because of something structural in the market, not because your RSI period is exactly 14 rather than 13 or 15. If your strategy needs exact numbers to perform, those numbers were fitted to the data you had. They are not evidence of anything repeatable.
We move each parameter by ±20% of the value you reported and compute the Sharpe ratio at each point. The sensitivity score is the coefficient of variation of those Sharpes, which is a plain measure of how much they bounce around. A flat plateau scores well. A sharp peak with steep sides triggers a warning.
The chart shows Sharpe against RSI period for a strategy fitted to RSI 14. A sharp peak at 14, with Sharpe dropping 78% at 12 or 16, is the classic overfitting shape. Strategies that look like this almost never survive live trading.
Sensitivity analysis following Ioannidis (2005) principles for stability checking. Parameter space exploration using ±20% variation with Sharpe ratio as the objective function.
In August 2026 we ran these eight tests on six real strategies whose quality was already known: a live, validated edge, its retired and degraded variants, and two strategies the research program that built them had killed. The engine was not told which was which.
The killed strategies graded F. The live edge graded B and beat everything that had been killed. Sorting all six by the engine's continuous outputs, the Monte Carlo p-value and the Probabilistic Sharpe Ratio, reproduces the known quality ordering exactly. Every audit finished in under 30 seconds. The report also lists what the engine gets wrong, because a methodology page that hides its limits is not a methodology page. Read the full calibration report →
Most strategies fail these tests. Yours probably will too.
Better to know now, while it's free, than after you've funded an account.
Start free, no credit card