Test Results
With 25 configurations tested, the minimum backtest length needed to stand behind a Sharpe ratio of 2.08 is about 179 trades. This backtest has 47. That is 26% of what the claim requires.
Insufficient sample: 47 trades available, ~179 required to validate the observed Sharpe ratio at this confidence level. Approximately 132 more trades needed before these results are trustworthy.
The overfitting test splits your configurations into a training half and a testing half, 12,870 different ways. Each split asks whether the in-sample winner also beat the median out-of-sample. A score above 0.5 means your winner is more likely than not a lucky draw from noise.
Moderate overfitting risk (54%). The results are statistically ambiguous, so the strategy may or may not have a genuine edge.
Score interpretation: <0.4 = low overfitting risk, 0.4–0.6 = moderate, >0.6 = high risk. Source: Bailey et al., “The Probability of Backtest Overfitting” (2015).
The trade log is split by date. The first 70% is in-sample, the last 30% is out-of-sample. This strategy kept 53% of its in-sample Sharpe ratio on the part it had never seen. A strategy that holds up keeps 70–80%. This one is in between. Some of the performance carried over to unseen data. A real part of it did not.
Moderate degradation: OOS Sharpe (1.38) is 53% of IS Sharpe (2.60). Some parameter overfitting may be present.
18.7% of randomly shuffled trade sequences with the same volatility matched or beat this Sharpe ratio. The strategy cannot be told apart from random chance. That is the single most disqualifying result a backtest can produce.
No statistically significant edge (p=0.1870). 18.7% of random permutations achieve similar or better performance. Results are consistent with random chance.
68% of the profits came from the best quarter of time periods. The HHI of 0.565 is past the 0.5 threshold. The strategy only worked in one kind of market.
High regime concentration: top 25% of periods contain 68% of profits (HHI=0.565). Only 50% of periods profitable. Strategy may fail outside its favoured regime.
After correcting for 25 reported configurations tested, the bar to beat is a Sharpe of 2.48. That is the best result you would expect from that many random strategies. The observed Sharpe of 2.08 does not clear it convincingly. A PSR of 43.7% leaves a 56.3% chance there is no edge here at all.
Not statistically significant after correcting for multiple testing (PSR=43.70%). Observed SR 2.08 does not significantly exceed the 25-trial benchmark SR of 2.48. This result could easily arise from selecting the best of many strategies.
Across the 20 tested parameter configurations, performance moves with a coefficient of variation of 0.346. That is moderate fragility. The strategy sits on a peak, and small changes to its main parameter walk it off that peak.
Moderate sensitivity (CV=0.35): some variation across 20 configs. The strategy has a performance peak but is not completely parameter-fragile.
Realistic futures execution costs (spread, commission, slippage) take 50.7% off the returns. This test is informational and carries no score penalty. A haircut this size is severe. For a strategy without a confirmed edge, it is disqualifying on its own.
Backtest Sharpe (idealized): 2.10. Realistic Sharpe (after costs): 1.04, a 51% degradation. Estimated annual cost drag: $212. Cost impact is severe for a futures strategy. The edge may not survive realistic execution.
OverfitCheck provides statistical analysis of backtest data only. Results do not constitute financial advice, investment recommendations, or a guarantee of future performance. Statistical robustness in historical testing does not predict live trading outcomes. All trading involves substantial risk of loss. The score and all test results are tools for your own informed decision-making. You are solely responsible for any trading decisions you make.
Test Results
Adequate sample: 581 trades. The minimum backtest length requirement is effectively zero at this trial count, so sample size is not a limiting factor.
Honesty note. At a reported trial count of 1, the minimum-backtest-length requirement works out to zero, so this test can only fail when more than one tested configuration is reported. The count still matters. The same number feeds the Deflated Sharpe correction, where it bites hard.
One configuration was uploaded, so the full Combinatorially Symmetric Cross-Validation could not run. It needs several configurations to compare against each other. A bootstrap proxy ran instead: the share of resampled trade sequences whose Sharpe matches or beats the observed one.
When this test can't tell you anything, we say so. The bootstrap distribution recenters on the observed Sharpe, so this proxy lands near 0.5 for any strategy at all. In our ground-truth calibration, all six real strategies scored between 0.467 and 0.501 on it, the live edge and the killed ones alike. Read this number as an advisory, not as evidence of overfitting. Upload several parameter configurations and the full 12,870-split CSCV algorithm replaces it.
Methodology and the full disclosure: see PBO & CSCV and the calibration write-up. Source: Bailey et al., “The Probability of Backtest Overfitting” (2015).
The trade log is split by date. The first 70% is in-sample, the last 30% is out-of-sample. This strategy kept 55% of its in-sample Sharpe ratio on the part it had never seen. A strategy that holds up keeps 70–80%. This one is in between. Some of the performance carried over to unseen data. A real part of it did not.
Moderate degradation: OOS Sharpe (0.89) is 55% of IS Sharpe (1.63). Some parameter overfitting may be present.
1.6% of randomly shuffled trade sequences with the same volatility matched or beat this Sharpe ratio. The edge is statistically significant at the 5% level. Random orderings of these trades rarely do this well.
Statistically significant edge (p=0.0157). Fewer than 5% of random permutations match this performance. Strategy likely captures real inefficiency.
75% of the profits came from the best quarter of time periods. Spread across periods (HHI 0.061) is not extreme, but a top-quarter share above 70% takes the full regime penalty. The profits arrived in bursts. Live, you would spend most of your time waiting between them.
High regime concentration: top 25% of periods contain 75% of profits (HHI=0.061). Only 61% of periods profitable. Strategy may fail outside its favoured regime.
After correcting for 1 reported trial, the bar to beat is a Sharpe of 0.00. That is the best result you would expect from that many random strategies. The observed Sharpe of 1.42 clears it. A PSR of 98.8% leaves a 1.2% chance there is no edge here at all.
At one reported trial there is almost nothing to correct for, so this PSR mostly reflects how long the sample is and how lopsided the returns are. If more configurations were searched than were reported, this test is being too kind.
Statistically significant Sharpe ratio (PSR=98.81%). Observed SR of 1.42 exceeds the expected maximum from random selection (0.00). Performance is unlikely due to chance.
This test needs several backtests at different parameter values. Only one was uploaded, so it did not run. No penalty applied.
Realistic futures execution costs (spread, commission, slippage) take 5.1% off the returns. This test is informational and carries no score penalty. At this size, a genuine edge can absorb the drag.
Backtest Sharpe (idealized): 1.42. Realistic Sharpe (after costs): 1.35, a 5% degradation. Estimated annual cost drag: $2,614. Transaction costs barely move this strategy.
OverfitCheck provides statistical analysis of backtest data only. Results do not constitute financial advice, investment recommendations, or a guarantee of future performance. Statistical robustness in historical testing does not predict live trading outcomes. All trading involves substantial risk of loss. The score and all test results are tools for your own informed decision-making. You are solely responsible for any trading decisions you make.