The question
"My backtest shows +17% ROI. Why shouldn't I believe it?"
Because you probably didn't run one backtest. You ran dozens, kept tweaking, and stopped when one looked good. This is why betting backtests lie, and it catches almost everyone who builds systems.
The idea in one sentence
When you test many ideas on the same data, the best-looking one is mostly the luckiest one, so the more you try, the higher the bar a result must clear before you believe it.
The picture
Press the button below. It creates 100 betting systems with no edge whatsoever: each is 200 bets at evens, where every bet is a fair coin toss. Then it plots the ROI of all 100 in a histogram.
Most land near zero, as you'd expect. But the spread is wide, because 200 bets is a small sample. Somewhere on the right-hand edge sits one system showing a handsome profit. That is "your winning system": the one you would have written up, staked and told your friends about.
Press it again. The winner changes every time, but there is always one, and it's usually somewhere between +15% and +20%.
We checked this by running the experiment 20,000 times. The best of 100 worthless systems had a median ROI of +17% before commission (+15.8% after 2%). At least one system showed +10% or better in 99.99% of runs.
| Ideas tested | Typical ROI of the best one (200 bets at evens, no edge) |
|---|---|
| 1 | 0% |
| 10 | +11% |
| 20 | +13% |
| 100 | +17% |
| 1,000 | +22% |
Worked Betfair example
You're building an Over 2.5 goals system on prices near 2.0. You try 100 variations: different leagues, form filters, days of the week, price bands. Each has 200 qualifying bets in your data. (Illustrative, but every number below is exact for zero-edge systems.)
- The winner. The best variation won 117 of 200 at 2.0. Before commission that's (117 − 83) ÷ 200 = +17%. After 2% commission on winnings: (117 × 0.98 − 83) ÷ 200 = +15.8%.
- Test it on its own. At evens, one bet swings by £1 per £1 staked, so the standard error over 200 bets is 1 ÷ √200 = 7.1%. That gives t = 17 ÷ 7.1 = 2.4. The exact chance of 117+ wins from 200 fair coin tosses is 0.0097, under 1%. On its own, that looks like strong evidence.
- Now remember the other 99. The chance that at least one of 100 worthless systems does this well is 1 − (1 − 0.0097)¹⁰⁰ = 62%. More likely than not.
- Count the fakes at the usual bar. With p below 0.05 as the test (113+ wins), you'd expect about 3.8 of the 100 worthless systems to pass, and at least one to pass 98% of the time.
- Raise the bar to match. The simplest correction, Bonferroni, divides the bar by the number of tests: 0.05 ÷ 100 = 0.0005. You'd now need 124 wins from 200, a +24% ROI (+22.8% after commission). Your +17% winner doesn't come close.
- Test it on fresh data. Run the winning rule on the next 200 matches it has never seen. If it has no edge, you should expect about 0% before commission and −1% after.
Verdict: a +17% backtest chosen from 100 tries is exactly what no edge looks like. It isn't evidence of anything until it survives step 6.
Overfitting is the same trap inside a model
Multiple testing is picking the best of many systems. Overfitting is the same thing inside one model: every extra rule, parameter or feature is another chance to fit the noise. A model with enough knobs can explain any history perfectly and predict nothing.
The warning signs:
- Amazing backtest, poor live results.
- Very specific rules, such as "away teams, Tuesdays, odds 2.2 to 2.6, after a draw".
- Fragile results. Move a date or a price band slightly and the profit disappears.
- No reason. You can't say who is on the wrong side of the bet, or why.
The formula
The chance of at least one fake winner
- α is the p-value bar for one test, e.g. 0.05.
- m is the number of independent ideas you tested.
In plain English: test 20 worthless ideas at the 5% bar and there's a 64% chance at least one passes. Test 100 and it's over 99%.
The Bonferroni correction
- α is the overall false-alarm rate you're willing to accept across all your tests.
- m is the number of tests.
In plain English: if you tried 100 ideas, each one has to be 100 times more convincing than a single planned test. It's strict when your ideas are close variations of each other, and gentler methods exist (multiple testing corrections), but it's a safe place to start.
How lucky the best of m can look
- Φ⁻¹ is the inverse of the normal curve: it turns a probability into a number of standard errors.
- m is the number of systems tested.
- σ is the swing of one bet (£1 at evens), and n is the number of bets.
In plain English: for 100 systems, z is about 2.46, and 2.46 × 7.1% ≈ +17%. That's the typical best result when nothing works, and the bar your real result has to beat by a distance.
Try it
Press "Generate 100 random systems" five times and note the winner each time. Then imagine you only ever saw the winner: that's what a backtest report shows you.
Common mistakes
- Only counting the tests you wrote down. Every tweak, league, filter and price band you looked at counts, including the ones you abandoned after a glance.
- Re-using the test data. Once you've looked at a hold-out period and changed the system because of it, it isn't a hold-out any more (Lesson 7.6).
- Adding rules until the losers disappear. Every rule that removes a few past losers is fitting noise unless it has a reason that would hold in future.
- Reading +17% as "at least some edge". The best of 100 is biased upwards by luck. Expect it to shrink a lot on fresh data, often to nothing.
- Ignoring the price. Many "systems" are just the market's own information rediscovered. Check the rule beats the closing price, not just the results (Lesson 2.4).
This is pitfall 5, the biggest killer of all: why good models still lose money.