In one sentence
Power analysis tells you how many bets you need so that, if you really have an edge of a given size, a statistical test will have a good chance of detecting it.
How it works
Every test can go wrong in two ways. It can "find" an edge that is not there (a false positive), or miss an edge that is (a false negative). Power is the chance of not missing a real edge, and 80% is the usual target.
Three things decide how many bets you need: the size of the edge, the odds you bet at, and how strict you are about false positives. Small edges and long odds need far more bets, because the natural swings are large compared with the edge.
The value of doing this first is that it sets expectations. Many bettors abandon a good strategy after 300 bets, or scale up a lucky one, when power analysis would have told them 300 bets could not answer the question either way.
The maths
- n: the number of level-stakes bets needed.
- z alpha: 1.645 for a one-sided test at the 5% level.
- z beta: 0.842 for 80% power.
- σ: the standard deviation of profit per bet, in stakes.
- ROI: the true edge you want to be able to detect, for example 0.05.
- p: the win rate implied by that ROI at odds O.
In plain English: the bets you need grow with the square of the noise-to-edge ratio, so halving the edge roughly quadruples the sample.
Worked betting example
You want to detect a +5% ROI, after commission, on Betfair football bets at average odds of 3.0, one-sided 5% test, 80% power.
- Win rate: (1 + 0.05) ÷ 3.0 = 35%.
- Spread per bet: σ ≈ 1.431 stakes.
- Noise-to-edge ratio: 1.431 ÷ 0.05 ≈ 28.6.
- Multiply by 1.645 + 0.842 = 2.487, giving about 71.2.
- Square it: n ≈ 5,064 bets.
That is more than the 3,276 bets at which a true 5% edge produces t = 2 on average. "On average" is the catch: luck pushes the real result above or below that, so a sample sized to hit the bar on average falls short of it about half the time. Power analysis sizes the sample so the expected t is about 2.49, and you clear the 1.645 bar 80% of the time.
How the answer changes (same 5% test, 80% power):
| Average odds | True ROI | Bets needed |
|---|---|---|
| 1.5 | +5% | ≈1,169 |
| 3.0 | +10% | ≈1,292 |
| 3.0 | +5% | ≈5,064 |
| 6.0 | +5% | ≈12,854 |
| 3.0 | +2% | ≈31,216 |
A realistic edge of 2% at odds of 3.0 needs over 31,000 bets to confirm with this confidence. At £10 a bet, that is over £310,000 turned over.
Where it's good
- Planning a trial of a new system: decide the sample size and stake size before the first bet.
- Judging tipsters: a 200-bet record on 8.0 shots is almost meaningless, and power analysis shows why.
- Choosing where to look: short-priced, high-volume markets can be tested far faster.
Limitations and pitfalls
- You must guess the edge size in advance. Bettors routinely assume bigger edges than they have, which makes the required sample look deceptively small.
- The formula assumes level stakes at roughly constant odds. Mixed odds need simulation, which is easy with Monte Carlo methods.
- Correlated bets, such as several in the same match, reduce the effective sample size, so you need more bets than calculated.
- Power analysis does not account for the edge fading while you test, which is common in efficient markets.
- Commission lowers your edge. Use the ROI after commission, not before.
- Low power combined with lucky early results leads to overestimating the edge, a trap known as the winner's curse.
How to build it
- Python:
statsmodels.stats.powerfor standard tests, or a short numpy simulation for mixed odds and staking. - Data: your expected odds distribution and a sober estimate of the edge after commission.
- Tip: simulate 10,000 seasons of your planned betting with the assumed edge and count how often you would pass your test; it doubles as a reality check on drawdowns.
Related methods
- Hypothesis testing - the test whose power you are planning for.
- P-values and confidence intervals - what you will report once the bets are in.
- Sequential testing (SPRT) - can reach a decision with fewer bets on average.
- Law of large numbers - why big samples reveal the true edge.
- Risk of ruin - making sure your bank survives the sample you need.