Blog/Cost forensics · Aug 3, 2026

The bad years were a toll booth

OverfitCheck Research · 6 min read

Every strategy with a long backtest has bad years, and every trader has a story about them. The usual story is regime. The market changed, the edge went dormant, low volatility starved the setups. Our story was the same one. An intraday breakout system on a Nasdaq-100 CFD, fourteen years of history, 867 trades, profitable overall, with an ugly cluster of losing years early in the sample. We believed the regime story for a long time. Then we audited it.

The audit was a forensic decomposition. Rebuild each losing year twice, once at the venue’s measured transaction costs and once at zero cost, then attribute the loss line by line. The result was not subtle. 91% of the aggregate loss in the bad years was spread cost. About 9% was genuine signal drought, the strategy’s actual edge failing to show up. At zero cost, 2013 and 2018 would have been ordinary drawdown years. With costs, they were disasters.

Venue toll
91%
Genuine drought
9%
Share of aggregate losing-year losses, decomposed at measured venue costs vs zero cost.

The geometry of a toll

The mechanism is arithmetic, not mystery. On a CFD you pay the spread on every round trip, and the spread is roughly constant in index points. Our strategy’s stop distance, by contrast, scales with volatility. Quiet markets mean tight stops. Define the cost ratio:

c  = spread ÷ stop distance
p* = (1 + c) ÷ (1 + R)

p* = breakeven win rate for a strategy risking 1 to make R.
At c = 0, breakeven is 1/(1+R). Every point of c raises the bar.

When volatility collapses, the stop collapses with it, the spread does not, and c balloons. The same fixed toll that is a rounding error against a wide stop becomes a meaningful fraction of a narrow one. Your entry logic has not changed. Your win rate has not changed. Your breakeven win rate has. If that arithmetic is new, lesson two of our course works it through with round numbers: costs are the first filter.

We measured the spread ourselves rather than trusting advertised figures. Over weeks of in-terminal capture at a major CFD broker, the median spread in our trading window was 2.30 index points. It was 2.50 in the pre-open and 0.80 near the close. A threefold difference by time of day, which means a cost model with one number in it is already wrong.

2013 was structurally unpayable

Now the punchline. In 2013, the deadest volatility year in the sample, stop distances compressed so far that the cost-loaded breakeven win rate reached 28.4%. The strategy’s realized win rate across the entire fourteen years, good years included, is 27.3%.

Read that as a mechanism, not a statistic. In 2013’s cost geometry, this strategy would have needed to trade better than it had ever traded over its whole life just to break even. The edge did not need to weaken to produce a losing year. A losing year was guaranteed by construction. The strategy was not failing. It was paying a toll that had quietly risen above the fare it could ever collect.

This reframing matters because the two diagnoses have opposite remedies. If the bad years were regime, you want a smarter filter, and you will probably overfit one. If the bad years were a toll, you want to stop trading when the toll is unpayable, which is a single inequality you can compute at entry time, before the trade, from the spread and the stop. Anyone running this exercise on their own record has a third possibility to rule out before either of these, which is that nothing was wrong at all: normal looks like failure.

What the fix is worth, honestly

We tested exactly that rule. Skip any trade whose stop distance is below 12× the measured median spread. Applied across the full history:

$20.3k → $6.5k
Max drawdown
21 → 12
Longest losing streak, trades
+$8.5k
Net P&L delta · not citable as alpha

The honest framing clause we attached to that last number is binding in our own research registry. The +$8.5k may never be cited as expected profit. 78% of the gross benefit sits in 2013 to 2019, and the rule gives back $17.9k across 2015, 2016 and 2020 by skipping trades that would have won. What the rule actually buys is the removal of the structurally unpayable class, which is insurance against a small-stop regime returning, plus a two-thirds cut in drawdown. That is worth having. It is not the same thing as more money, and a backtest delta that cannot tell the difference is how cost rules get oversold.

One implementation detail turned out to matter more than we expected. The floor is registered as a formula, a fixed multiplier times the currently measured spread, never as a naked constant. When our measured spread basis changed from 2.0 to 2.30 points, the floor moved with it mechanically. The alternative, picking whichever historical constant makes the backtest look best, is cost-model curve fitting with extra steps, and it is banned outright in our research registry.

The population-scale echo

We treated the 91% figure with suspicion at first. It seemed too clean. Then we noticed it is not unusual at all. Barber, Lee, Liu and Odean’s complete-market study of Taiwanese day trading (Review of Financial Studies, 2009) found that roughly two-thirds of aggregate individual-investor losses are frictions: commissions, transaction tax, spreads, rather than bad predictions. Recent work on retail zero-days-to-expiry option flow attributes roughly 70% of retail losses to transaction costs. The retail population’s bad years are mostly a toll booth too. Ours just came with a trade log detailed enough to prove it.

What to do with your own backtest

Run it twice. Once at zero cost, once at costs you measured yourself, at the time of day you actually trade. Then attribute the difference year by year. Compute your breakeven win rate per year at that year’s stop geometry and plot it against your realized win rate. Any year where the lines cross was not a drawdown. It was a toll you could not pay, and no entry filter, regime story, or added indicator will fix it. This arithmetic is exactly what OverfitCheck’s transaction-cost test runs on every uploaded backtest. Your equity curve recomputed under realistic venue costs, with a flag for the periods where the breakeven math quietly moved out of reach.

Read next

Run your backtest through eight peer-reviewed statistical tests. Free, no credit card required.

Try OverfitCheck free →