Course·Lesson 3 of 5

How many trades you actually need

There is no single number, but there is a way to think about it that beats guessing.

A backtest with 30 trades tells you close to nothing. Not because 30 is a magic threshold, but because the range of outcomes that 30 trades is consistent with is enormous.

Take a coin, which has no edge at all. Flip it 30 times and getting 18 heads is unremarkable. That is a 60 percent win rate produced by a fair coin. If those had been trades you would be looking at a 60 percent strategy and feeling confident. Flip the same coin 1,000 times and 600 heads essentially never happens. The larger sample squeezes out the luck.

This is the whole idea. Sample size does not make a strategy better. It narrows the range of things the result could be hiding.

Why the honest answer depends on your search

Here is where the previous lesson comes back. The number of trades you need depends on how many variations you tried, because searching harder makes impressive results cheaper to obtain by luck alone.

If you tested one idea and it worked, a moderate sample is meaningful evidence. If you tested four thousand combinations and kept the best, you need substantially more evidence to make the same claim, because you have given randomness four thousand chances to produce something that looks good.

This is the part people find backwards, so it is worth stating plainly: more searching does not get you to a conclusion faster. It raises the bar you have to clear.

The uncomfortable one

A bigger claimed return needs more evidence, not less. A strategy claiming an extraordinary return from a short record is not more impressive than a modest one. It is less believable, because extraordinary results are exactly what a small sample produces by accident.

Trades, not days

Count trades, not calendar time. Ten years of history containing 40 trades is a 40 trade sample. The decade sounds reassuring and contributes almost nothing, because the thing being estimated is the behaviour of a trade and you have 40 examples of it.

This catches people running strategies on higher timeframes, where a long history produces very few events. A weekly-bar system tested over fifteen years may have fewer trades than a day-trading system has in a month, and it deserves correspondingly less confidence, however long the chart looks.

A sample can also shrink after the fact, when part of it turns out not to be what you thought. A broker clock error once put half of a backtest’s trades in a different session than every analysis assumed, which left the claim resting on far fewer usable trades than the trade count suggested: the clock that lied for ten years.

What to do about a small sample

The honest options are limited, and none of them involve reinterpreting the numbers you have.

If you would rather have this computed than estimated, the minimum sample size test works out how many trades your reported result needs behind it, given how many configurations you tested.

Next: why long flat stretches are normal, and why that is dangerous to know without also knowing what a real breakdown looks like.