Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Module 7 · Lesson 7.4

Overfitting and multiple testing

“Why do backtests lie?”

The question

"My backtest shows +17% ROI. Why shouldn't I believe it?"

Because you probably didn't run one backtest. You ran dozens, kept tweaking, and stopped when one looked good. This is why betting backtests lie, and it catches almost everyone who builds systems.

The idea in one sentence

When you test many ideas on the same data, the best-looking one is mostly the luckiest one, so the more you try, the higher the bar a result must clear before you believe it.

The picture

Press the button below. It creates 100 betting systems with no edge whatsoever: each is 200 bets at evens, where every bet is a fair coin toss. Then it plots the ROI of all 100 in a histogram.

Most land near zero, as you'd expect. But the spread is wide, because 200 bets is a small sample. Somewhere on the right-hand edge sits one system showing a handsome profit. That is "your winning system": the one you would have written up, staked and told your friends about.

Press it again. The winner changes every time, but there is always one, and it's usually somewhere between +15% and +20%.

Try it · Random system generator
Your ‘winning’ system: +15.8%-20%-10%0%10%20%30%
Best system
+15.8%
Median system
-1.0%
Worst system
-13.9%
+10% or better
5 of 100
Not one of these systems has an edge, yet the best shows +15.8% over 200 bets. Over its next 200 bets it should expect about -1.0%, the same as the rest.
Winner each time you pressed (1)
+15.8%
Every system here is a fair coin: 200 bets at evens with no edge whatsoever. Each square is one system, placed by its ROI.

We checked this by running the experiment 20,000 times. The best of 100 worthless systems had a median ROI of +17% before commission (+15.8% after 2%). At least one system showed +10% or better in 99.99% of runs.

Ideas tested Typical ROI of the best one (200 bets at evens, no edge)
1 0%
10 +11%
20 +13%
100 +17%
1,000 +22%

Worked Betfair example

You're building an Over 2.5 goals system on prices near 2.0. You try 100 variations: different leagues, form filters, days of the week, price bands. Each has 200 qualifying bets in your data. (Illustrative, but every number below is exact for zero-edge systems.)

  1. The winner. The best variation won 117 of 200 at 2.0. Before commission that's (117 − 83) ÷ 200 = +17%. After 2% commission on winnings: (117 × 0.98 − 83) ÷ 200 = +15.8%.
  2. Test it on its own. At evens, one bet swings by £1 per £1 staked, so the standard error over 200 bets is 1 ÷ √200 = 7.1%. That gives t = 17 ÷ 7.1 = 2.4. The exact chance of 117+ wins from 200 fair coin tosses is 0.0097, under 1%. On its own, that looks like strong evidence.
  3. Now remember the other 99. The chance that at least one of 100 worthless systems does this well is 1 − (1 − 0.0097)¹⁰⁰ = 62%. More likely than not.
  4. Count the fakes at the usual bar. With p below 0.05 as the test (113+ wins), you'd expect about 3.8 of the 100 worthless systems to pass, and at least one to pass 98% of the time.
  5. Raise the bar to match. The simplest correction, Bonferroni, divides the bar by the number of tests: 0.05 ÷ 100 = 0.0005. You'd now need 124 wins from 200, a +24% ROI (+22.8% after commission). Your +17% winner doesn't come close.
  6. Test it on fresh data. Run the winning rule on the next 200 matches it has never seen. If it has no edge, you should expect about 0% before commission and −1% after.

Verdict: a +17% backtest chosen from 100 tries is exactly what no edge looks like. It isn't evidence of anything until it survives step 6.

Overfitting is the same trap inside a model

Multiple testing is picking the best of many systems. Overfitting is the same thing inside one model: every extra rule, parameter or feature is another chance to fit the noise. A model with enough knobs can explain any history perfectly and predict nothing.

The warning signs:

  • Amazing backtest, poor live results.
  • Very specific rules, such as "away teams, Tuesdays, odds 2.2 to 2.6, after a draw".
  • Fragile results. Move a date or a price band slightly and the profit disappears.
  • No reason. You can't say who is on the wrong side of the bet, or why.

The formula

The chance of at least one fake winner

P(at least one false positive)=1−(1−α)mP(\text{at least one false positive}) = 1 - (1 - \alpha)^m
  • α is the p-value bar for one test, e.g. 0.05.
  • m is the number of independent ideas you tested.

In plain English: test 20 worthless ideas at the 5% bar and there's a 64% chance at least one passes. Test 100 and it's over 99%.

The Bonferroni correction

αeach=αm\alpha_{\text{each}} = \frac{\alpha}{m}
  • α is the overall false-alarm rate you're willing to accept across all your tests.
  • m is the number of tests.

In plain English: if you tried 100 ideas, each one has to be 100 times more convincing than a single planned test. It's strict when your ideas are close variations of each other, and gentler methods exist (multiple testing corrections), but it's a safe place to start.

How lucky the best of m can look

zbest≈Φ−1(0.51/m),ROIbest≈zbest×σnz_{\text{best}} \approx \Phi^{-1}\big(0.5^{1/m}\big), \qquad \text{ROI}_{\text{best}} \approx z_{\text{best}} \times \frac{\sigma}{\sqrt{n}}
  • Φ⁻¹ is the inverse of the normal curve: it turns a probability into a number of standard errors.
  • m is the number of systems tested.
  • σ is the swing of one bet (£1 at evens), and n is the number of bets.

In plain English: for 100 systems, z is about 2.46, and 2.46 × 7.1% ≈ +17%. That's the typical best result when nothing works, and the bar your real result has to beat by a distance.

Try it

Press "Generate 100 random systems" five times and note the winner each time. Then imagine you only ever saw the winner: that's what a backtest report shows you.

Common mistakes

  • Only counting the tests you wrote down. Every tweak, league, filter and price band you looked at counts, including the ones you abandoned after a glance.
  • Re-using the test data. Once you've looked at a hold-out period and changed the system because of it, it isn't a hold-out any more (Lesson 7.6).
  • Adding rules until the losers disappear. Every rule that removes a few past losers is fitting noise unless it has a reason that would hold in future.
  • Reading +17% as "at least some edge". The best of 100 is biased upwards by luck. Expect it to shrink a lot on fresh data, often to nothing.
  • Ignoring the price. Many "systems" are just the market's own information rediscovered. Check the rule beats the closing price, not just the results (Lesson 2.4).

This is pitfall 5, the biggest killer of all: why good models still lose money.

Check yourself

1. You test 100 variations of a system and the best has a p-value of 0.01. What's the problem?
2. Which of these is the strongest warning sign of an overfitted system?
3. Your best system showed +17% over 200 bets. What should you expect over its next 200 bets if it has no real edge?
Key takeaway

The best of many backtests is always flattered by luck. Count every idea you tried, raise the bar to match, and only believe what survives on data the system never saw.

Go deeper in the Model Library
Next lesson
7.5 Building an honest backtest →
What must my backtest include?
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
MembersWhy Betting Backtests Lie: Overfitting and Multiple Testing Explained — Statometrics Academy