In one sentence
A p-value tells you how often pure luck would produce results at least as good as yours, and a confidence interval gives a range of true ROIs that are consistent with your record.
How it works
The p-value answers a narrow question: if I had no edge at all, how often would I see a record this good or better? A small p-value, such as 0.02, means luck rarely produces this. It does not tell you the chance that you have an edge, a common and costly misreading.
A confidence interval is often more useful. Instead of one yes-or-no verdict, it gives a range, such as "my true ROI is plausibly between −4% and +14%". If the range is wide and includes zero, your record has not yet separated skill from luck.
Both come from the same ingredient: the standard error, which measures how much your ROI would wobble from luck alone over that many bets.
The maths
- ROI with a bar: your observed average profit per £1 staked.
- σ: the standard deviation of profit on a single bet, in stakes.
- n: the number of bets.
- 1.96: the multiplier that covers 95% of a normal distribution.
- Z: a standard normal variable, used because averages of many bets are close to normally distributed.
In plain English: your true ROI is probably within about two standard errors of what you observed, and the p-value is the chance of beating your result with no edge.
Worked betting example
You have 1,000 level-stakes Betfair bets, draws and away wins at average odds of 3.0, with a +5% ROI after commission, so £500 profit at £10 stakes and a 35% win rate.
- Single-bet standard deviation: about 1.431 stakes.
- Standard error: 1.431 ÷ √1,000 ≈ 0.0452, or 4.52%.
- t statistic: 0.05 ÷ 0.0452 ≈ 1.10.
- One-sided p-value: about 0.135. A no-edge bettor matches or beats this roughly 1 time in 7.
- 95% confidence interval: 5% ± 1.96 × 4.52%, so roughly −3.9% to +13.9%.
That interval is the key message. Your record is consistent with a losing strategy and with an excellent one. You cannot tell which yet.
Now the same +5% ROI over 3,000 bets:
- Standard error falls to about 2.61%.
- t ≈ 1.91, one-sided p ≈ 0.028.
- 95% interval: about −0.1% to +10.1%.
Notice the one-sided p-value is below 0.05, but the two-sided 95% interval still just touches zero. They are asking slightly different questions; neither is wrong. About 3,276 bets would be needed to reach t = 2.
Where it's good
- Reporting results fairly, to yourself or to subscribers, with a range rather than a single headline ROI.
- Deciding stake sizes: planning around the lower end of the interval is a sensible safeguard.
- Comparing two strategies, by looking at whether their intervals overlap heavily.
- Spotting tipsters whose records are too short to judge, however good they look.
Limitations and pitfalls
- A p-value is not the probability that you have no edge. It is the probability of your data assuming no edge.
- The 0.05 line is a convention, not a law. A p of 0.049 and a p of 0.051 are practically the same evidence.
- Testing many strategies and reporting the one with the smallest p-value makes it meaningless; correct for multiple testing.
- Checking the p-value after every bet and stopping when it dips below 0.05 greatly raises the false positive rate.
- The normal approximation weakens with long odds and small samples; bootstrap the interval instead.
- Correlated bets, such as several in one match, make the real interval wider than the formula shows.
- Commission must already be in your profit figures, or the interval is centred in the wrong place.
How to build it
- Python:
scipy.statsfor the normal and t distributions;numpyfor the standard error;scipy.stats.bootstrapfor a resampled interval. - Data: the full list of per-bet profits after commission.
- Tip: show the interval on a chart that updates as bets are added; watching it narrow teaches patience better than any p-value.
Related methods
- Hypothesis testing - the framework the p-value belongs to.
- Power analysis - how many bets you need for a narrow interval.
- Bootstrapping - intervals without the normal assumption.
- Multiple testing corrections - keeping p-values honest when testing many ideas.
- Central limit theorem - why the normal approximation works for many bets.