Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Evaluation, testing and behaviour

Multiple Testing Corrections

Adjusts for the fact that if you test enough betting systems, some will look profitable by pure luck.

Intermediateevaluation

In one sentence

Multiple testing corrections raise the bar for "significant" when you test many ideas at once, so that lucky systems do not get mistaken for real edges.

How it works

Every test at the usual 5% level has a 1 in 20 chance of a false alarm when there is no edge. Test one system and that risk is small. Test 20 systems, filters or leagues and you should expect about one to look significant from luck alone.

This is how most "proven" betting systems are born. Someone scans thousands of angles, such as away favourites on Tuesdays in the Championship after a defeat, and the best-looking one is presented as a discovery. Its p-value is meaningless because nobody counted the other attempts.

Corrections fix this. Bonferroni divides the 5% threshold by the number of tests, which is safe but strict. Holm is a step-by-step version that is never less powerful. Benjamini-Hochberg controls the false discovery rate, meaning the expected share of your "discoveries" that are flukes, and is less strict again.

The maths

P(at least one false positive)=1−(1−α)mP(\text{at least one false positive}) = 1 - (1 - \alpha)^m Bonferroni: reject if pi≤αm,Benjamini-Hochberg: largest k with p(k)≤km α\text{Bonferroni: reject if } p_i \le \frac{\alpha}{m}, \qquad \text{Benjamini-Hochberg: largest } k \text{ with } p_{(k)} \le \frac{k}{m}\,\alpha
  • α: the significance level for a single test, usually 0.05.
  • m: the number of tests you ran.
  • p with subscript: the p-value of each test; p with a bracketed k is the k-th smallest.

In plain English: the more you test, the more likely a fluke; so either shrink the threshold for every test or judge them in order from strongest to weakest.

Worked betting example

With 20 independent systems and no real edges, the chance at least one clears p = 0.05 is 1 − 0.95 to the power 20 ≈ 64%. Bonferroni would demand p of 0.0025 or smaller from each.

Now a smaller case. You back-tested six football angles, such as Over 2.5 goals in particular leagues, and got these one-sided p-values, sorted smallest first: 0.004, 0.011, 0.020, 0.030, 0.200, 0.450.

Uncorrected: four angles pass 0.05.

Bonferroni: threshold 0.05 ÷ 6 ≈ 0.0083. Only 0.004 passes. One angle.

Holm: compare the smallest with 0.05 ÷ 6 ≈ 0.0083 (pass), the next with 0.05 ÷ 5 = 0.010. Since 0.011 is above 0.010, stop. One angle.

Benjamini-Hochberg: thresholds are 0.0083, 0.0167, 0.025, 0.0333, 0.0417, 0.05. The largest rank where the p-value is under its threshold is rank 4 (0.030 against 0.0333). So the first four pass, with the expected share of flukes among them held at or below 5% on average.

Which to use depends on the cost of a mistake. If each "discovery" means real money at real stakes, use Holm. If you are shortlisting ideas for further out-of-sample testing, Benjamini-Hochberg is reasonable.

Where it's good

  • Screening many system filters, leagues, bet types or price bands.
  • Judging a tipster service that promotes its best-performing tipster out of many.
  • Feature selection, when testing dozens of candidate variables for a model.
  • Any research session where you tried more than one thing, which is nearly all of them.

Limitations and pitfalls

  • You must count every test you ran, including the ones you abandoned. Most people forget, and the correction becomes too lenient.
  • Tests on overlapping data are correlated, so Bonferroni becomes overly strict; permutation-based methods handle this better.
  • Being strict means missing some real but small edges. There is no free lunch between false alarms and missed edges.
  • Corrections only deal with luck. They do not fix look-ahead bias, bad data or ignoring commission.
  • Many systems do not have a clean p-value to start with, especially when stakes vary or bets overlap.
  • The best defence is still a fresh out-of-sample period that you only look at once.

How to build it

  • Python: statsmodels.stats.multitest.multipletests with methods such as bonferroni, holm and fdr_bh.
  • Data: the p-value of every test you ran, logged automatically by your research code.
  • Tip: keep a research log that counts every idea you test; the number in that log is your m.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members