Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Evaluation, testing and behaviour

P-values and Confidence Intervals

A p-value measures how surprising your record is if you had no edge; a confidence interval gives a plausible range for your true ROI.

Beginnerevaluationstaking

In one sentence

A p-value tells you how often pure luck would produce results at least as good as yours, and a confidence interval gives a range of true ROIs that are consistent with your record.

How it works

The p-value answers a narrow question: if I had no edge at all, how often would I see a record this good or better? A small p-value, such as 0.02, means luck rarely produces this. It does not tell you the chance that you have an edge, a common and costly misreading.

A confidence interval is often more useful. Instead of one yes-or-no verdict, it gives a range, such as "my true ROI is plausibly between −4% and +14%". If the range is wide and includes zero, your record has not yet separated skill from luck.

Both come from the same ingredient: the standard error, which measures how much your ROI would wobble from luck alone over that many bets.

The maths

CI95%=ROI‾±1.96×σn\text{CI}_{95\%} = \overline{\text{ROI}} \pm 1.96 \times \frac{\sigma}{\sqrt{n}} p=P(Z≥ROI‾σ/n)p = P\left(Z \ge \frac{\overline{\text{ROI}}}{\sigma / \sqrt{n}}\right)
  • ROI with a bar: your observed average profit per £1 staked.
  • σ: the standard deviation of profit on a single bet, in stakes.
  • n: the number of bets.
  • 1.96: the multiplier that covers 95% of a normal distribution.
  • Z: a standard normal variable, used because averages of many bets are close to normally distributed.

In plain English: your true ROI is probably within about two standard errors of what you observed, and the p-value is the chance of beating your result with no edge.

Worked betting example

You have 1,000 level-stakes Betfair bets, draws and away wins at average odds of 3.0, with a +5% ROI after commission, so £500 profit at £10 stakes and a 35% win rate.

  1. Single-bet standard deviation: about 1.431 stakes.
  2. Standard error: 1.431 ÷ √1,000 ≈ 0.0452, or 4.52%.
  3. t statistic: 0.05 ÷ 0.0452 ≈ 1.10.
  4. One-sided p-value: about 0.135. A no-edge bettor matches or beats this roughly 1 time in 7.
  5. 95% confidence interval: 5% ± 1.96 × 4.52%, so roughly −3.9% to +13.9%.

That interval is the key message. Your record is consistent with a losing strategy and with an excellent one. You cannot tell which yet.

Now the same +5% ROI over 3,000 bets:

  • Standard error falls to about 2.61%.
  • t ≈ 1.91, one-sided p ≈ 0.028.
  • 95% interval: about −0.1% to +10.1%.

Notice the one-sided p-value is below 0.05, but the two-sided 95% interval still just touches zero. They are asking slightly different questions; neither is wrong. About 3,276 bets would be needed to reach t = 2.

Where it's good

  • Reporting results fairly, to yourself or to subscribers, with a range rather than a single headline ROI.
  • Deciding stake sizes: planning around the lower end of the interval is a sensible safeguard.
  • Comparing two strategies, by looking at whether their intervals overlap heavily.
  • Spotting tipsters whose records are too short to judge, however good they look.

Limitations and pitfalls

  • A p-value is not the probability that you have no edge. It is the probability of your data assuming no edge.
  • The 0.05 line is a convention, not a law. A p of 0.049 and a p of 0.051 are practically the same evidence.
  • Testing many strategies and reporting the one with the smallest p-value makes it meaningless; correct for multiple testing.
  • Checking the p-value after every bet and stopping when it dips below 0.05 greatly raises the false positive rate.
  • The normal approximation weakens with long odds and small samples; bootstrap the interval instead.
  • Correlated bets, such as several in one match, make the real interval wider than the formula shows.
  • Commission must already be in your profit figures, or the interval is centred in the wrong place.

How to build it

  • Python: scipy.stats for the normal and t distributions; numpy for the standard error; scipy.stats.bootstrap for a resampled interval.
  • Data: the full list of per-bet profits after commission.
  • Tip: show the interval on a chart that updates as bets are added; watching it narrow teaches patience better than any p-value.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members