We run a live, profitable intraday breakout strategy on a Nasdaq-100 CFD. Fourteen years of backtest, 867 trades, validated the slow way. Which means we know the itch every systematic trader knows. The working system that could surely be a little better. A trailing stop to protect open profit. A breakeven move so winners can’t become losers. Partial exits to smooth the equity curve. Every forum recommends them, and every backtesting platform makes them one checkbox away.
So we tested the whole toolkit properly. Eleven structural modifications, each one pre-registered, meaning the rule and its pass/fail threshold were frozen in writing before the first run. Each was then judged against the unmodified baseline on identical data. Ten of eleven made the strategy worse. The eleventh’s profit improvement was, on inspection, plausibly noise.
The battery
The eleven covered the standard advice. Trailing stops, a move-to-breakeven rule, partial profit-taking, a time-based exit, re-entry after a stop-out, an opposite-direction re-entry, a long-only variant, an entry-quality filter that blocked setups whose pre-entry range looked too wide and sloppy, and variations on those themes. Nothing exotic. This is the folk wisdom of trade management, applied to an edge we already knew was real.
| Modification class | What happened | Verdict |
|---|---|---|
| Trailing stop, breakeven move, partial exits | Net profit down in every variant. Each rule mostly truncated winners | DEAD |
| Entry-quality filter (block wide, sloppy setups) | Delivered 77% of baseline net. The ugly trades carried the profits | DEAD |
| Long-only variant | Directional advantage sign-flipped between eras | DEAD |
| Re-entry rules (same and opposite direction) | Negative expectancy added on top of the base system | DEAD |
| Time-based exit | Profit delta positive, bootstrap p = 0.23; risk reduction real | DEAD (kept off) |
The pattern behind the failures is structural, not bad luck. This strategy wins roughly 27% of the time at a large reward multiple. The entire edge lives in the right tail of the winners. Every trade-management overlay works by touching trades mid-flight, and what it mostly touches is that tail. A trailing stop turns eventual full-size winners into small ones. A breakeven move turns them into scratches. Partials cap the exact trades that pay for the 73% that lose. Each rule is insurance priced above what it is worth, sold to the part of you that dislikes watching open profit breathe.
The entry-quality filter deserves its own paragraph, because it is the most intuitive of the lot. Block the trades that look bad, keep the ones that look clean. Its result, 77% of baseline profit, is the general finding in miniature. The backtest’s ugliest, widest, least comfortable setups were where the money was.
The survivor, examined
One modification passed its frozen threshold. A time-based exit that closes the trade if it has not resolved within a fixed window. Its backtest profit delta was positive. Then we did the thing the checkbox never does. We bootstrapped it. Resampling daily P&L 10,000 times put the probability of seeing a profit improvement at least that large, under the null of no true improvement, at p = 0.23. Roughly one in four. That is not evidence. That is a coin leaning slightly on its edge.
Its risk reduction, on the other hand, was real. Consistent across resamples, and mechanically unsurprising, because less time in the market is less exposure. So the honest accounting reads: profit claim unproven, risk claim true. Under our rules that is not an improvement. It is a risk lever, documented and left switched off. Even the winner didn’t ship.
The trap that makes 11 trials dangerous
Suppose all eleven modifications were worthless. Pure noise. Selecting the best-performing one still guarantees you something that looks good. The expected maximum of N independent null results grows like √(2 ln N), so the best of eleven worthless ideas is expected to sit around two standard deviations above zero. The improvement you found by searching is indistinguishable from the improvement the search itself manufactures, unless you explicitly correct for the number of trials. The plain-English version of that mechanism, without the algebra, is lesson one of our short course: what overfitting actually is.
A second episode from the same research program makes the point sharper. In a separate pass we ran roughly 25 conditioning splits over the same trade log. Day-of-week, gap direction, prior-day behavior, news days. Exactly one came back with a 95% confidence interval excluding zero. Thursdays, at +$185 per trade. Across 25 comparisons, the expected number of false positives at that confidence level is about 1.25. The single strongest result in the entire exercise was precisely what noise predicts, so it was recorded as the batch’s designated multiplicity casualty and discarded. No mechanism, no trade.
None of this is our invention. Harvey, Liu and Zhu (2016) showed that after accounting for the profession’s collective data mining, a newly discovered factor should clear a t-statistic of 3, not 2. Bailey and López de Prado (2014) built the Deflated Sharpe Ratio, which discounts an observed Sharpe for the number of trials behind it, their variance, the sample length, and the shape of the return distribution. Bailey, Borwein, López de Prado and Zhu (2015) built the Probability of Backtest Overfitting to measure how often an in-sample winner loses its rank out of sample. The tools for pricing your own search effort have existed for a decade.
What actually protects you
Four habits did the work in our case. Freeze the rule and its threshold before the first run, so the result cannot negotiate with you. Count every trial. Our program’s registry currently records over sixty dead, individually logged hypotheses, and that count is an input to every new claim. Treat “no change” as the default winner, because on a genuine edge the base rate of improvements that survive honest testing is one in eleven on a good day. And require a mechanism, a reason the market must pay you, before a delta is allowed to become a rule, however pretty that delta looks. This is also why OverfitCheck’s upload form asks how many configurations you tested. The Probability of Backtest Overfitting and the Deflated Sharpe Ratio are computed against that number, and the verdict on your strategy honestly changes with it.