Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Evaluation, testing and behaviour

Hypothesis Testing

A structured way to ask whether a betting record shows real skill or could easily be luck, by testing it against a no-edge assumption.

Beginnerevaluationstaking

In one sentence

Hypothesis testing starts by assuming you have no edge, then asks how surprising your actual results would be if that were true.

How it works

You set up two competing claims. The null hypothesis says your true long-run ROI is zero: any profit is luck. The alternative says your true ROI is above zero. You then measure how far your results sit from zero in units of "normal luck", known as the standard error.

That distance is the test statistic, usually called t. A t of 1 means your profit is about one standard error of luck above zero, which happens all the time by chance. A t of 2 or more is the usual rough bar for "hard to explain by luck alone".

It works like a court case. The null hypothesis is "innocent" (no edge) until the evidence is strong enough. Failing to reject it does not prove you have no edge; it means the evidence is not yet strong enough.

The maths

t=ROI‾SE,SE=σnt = \frac{\overline{\text{ROI}}}{\text{SE}}, \qquad \text{SE} = \frac{\sigma}{\sqrt{n}} σ=p (O−1)2+(1−p)−ROI‾ 2\sigma = \sqrt{p\,(O - 1)^2 + (1 - p) - \overline{\text{ROI}}^{\,2}}
  • ROI with a bar: your average profit per £1 staked, for example 0.05 for +5%.
  • SE: the standard error, how much ROI naturally wobbles from luck over n bets.
  • σ: the standard deviation of profit on a single level-stakes bet, in stakes.
  • n: the number of bets.
  • p: your win rate; O: the average decimal odds.

In plain English: divide your ROI by the amount of wobble luck would produce over that many bets; the bigger the answer, the less likely it is luck.

Worked betting example

You have placed 1,000 level-stakes bets on Betfair, draws and away wins at average odds of 3.0, and made a +5% ROI after commission, which is £500 profit at £10 a bet.

  1. Win rate: a +5% ROI at odds 3.0 means p × 3.0 = 1.05, so p = 35%.
  2. Spread of a single bet: a winner returns +2 stakes, a loser −1. Plugging in gives σ ≈ 1.431 stakes.
  3. Standard error: 1.431 ÷ √1,000 ≈ 0.0452, or 4.52% of ROI.
  4. Test statistic: 0.05 ÷ 0.0452 ≈ 1.10.
  5. One-sided p-value: about 0.135.

So a bettor with no edge at all would produce a record this good or better roughly 13.5% of the time. That is not strong evidence.

Keep the same +5% ROI and extend to 3,000 bets. The standard error shrinks to about 0.0261, t rises to about 1.91 and the p-value falls to about 0.028. To reach t = 2 you need about 3,276 bets.

The lesson: at odds around 3.0, even a respectable 5% edge takes thousands of bets to show up clearly.

Where it's good

  • Judging your own record, a tipster's record, or a system you found online.
  • Deciding whether to scale up stakes or keep testing at small stakes.
  • Comparing a strategy with a benchmark, such as ROI versus the commission you pay.
  • Setting a rule in advance, such as "I will increase stakes only once t exceeds 2", which removes emotion from the decision.

Limitations and pitfalls

  • The normal approximation behind t works less well with very long odds and few bets. For long-odds bets such as Correct Score, use a permutation test or simulation instead.
  • Testing many systems and reporting the best is a trap: with 20 systems, one will look significant by luck. Use multiple testing corrections.
  • Stopping the moment the result looks significant inflates false positives. Decide the sample size in advance, or use sequential testing.
  • Level stakes are assumed. Variable staking changes the spread of results and the maths.
  • Bets that are linked, such as several bets in the same match, are not independent, so the real uncertainty is larger than the formula says.
  • Statistical significance is not the same as a useful edge. A tiny edge on a huge sample can be significant and still not worth the time after commission.
  • Past edge may not persist; markets adapt.

How to build it

  • Python: scipy.stats.ttest_1samp on the list of per-bet profits, with the alternative set to "greater".
  • Data: every bet placed, with stake, odds and settled profit after commission, not just a summary.
  • Tip: record your hypothesis and sample size before you start betting, and stick to them.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members