Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Evaluation, testing and behaviour

Calibration

Checks whether events you rate at 40% really happen about 40% of the time, the property that makes a model's probabilities safe to bet on.

Beginnerpre-matchin-playevaluationstaking

In one sentence

A model is well calibrated when its probabilities match reality on average: of all the selections it rates at 40%, roughly 40% win.

How it works

Group your past forecasts into buckets, say 0-10%, 10-20% and so on, then compare the average forecast in each bucket with how often those selections actually won. Plot one against the other; a perfectly calibrated model sits on the diagonal line. This plot is called a reliability diagram.

Calibration matters more for betting than almost any other property. You compare your probability with the odds to decide whether a bet has value, so if your 40% is really 31%, you will see value that is not there and bet into losses.

Calibration is not the same as being useful. A model that gives every match the league averages, say 45% home, 27% draw and 28% away, is well calibrated and tells you nothing. You want calibration and sharpness, meaning forecasts that are confident when they should be.

The maths

o^k=wins in bucket knk,zk=o^k−pˉkpˉk(1−pˉk)/nk\hat{o}_k = \frac{\text{wins in bucket } k}{n_k}, \qquad z_k = \frac{\hat{o}_k - \bar{p}_k}{\sqrt{\bar{p}_k (1 - \bar{p}_k) / n_k}}
  • ô: the observed win rate in bucket k.
  • n: the number of forecasts in bucket k.
  • p̄: the average forecast probability in bucket k.
  • z: how many standard errors the observed rate sits from the forecast; beyond about ±2 is a warning sign.

In plain English: count how often each bucket actually won, and check whether the gap from the forecast is bigger than luck would explain.

Worked betting example

Your Match Odds model has produced 470 home-win forecasts across three buckets (illustrative figures).

Bucket (avg forecast) Bets Winners Observed rate Standard error z
20% 150 33 22.0% 3.27% +0.61
40% 200 62 31.0% 3.46% −2.60
60% 120 70 58.3% 4.47% −0.37

Step by step for the 40% bucket: 62 ÷ 200 = 0.31. Standard error = √(0.4 × 0.6 ÷ 200) ≈ 0.0346. The gap is −0.09, which is −2.60 standard errors, unlikely to be bad luck.

Why it matters: say a team in that bucket is available at 2.8. Your model sees expected value of 0.40 × 2.8 − 1 = +12%. If the true rate is 31%, the real expected value is 0.31 × 2.8 − 1 = −13.2%, before commission. The overconfidence has turned an apparent value bet into a clear loser.

The 20% and 60% buckets are within normal noise, so the problem is local to mid-priced teams, and a recalibration step is needed.

Where it's good

  • A basic health check on any model before its probabilities go anywhere near a staking plan.
  • Spotting systematic bias, such as overrating favourites or underrating draws.
  • Checking the market itself: the favourite-longshot bias is a calibration failure of market prices.
  • Monitoring a live model over time to catch drift after rule changes or new seasons.

Limitations and pitfalls

  • Buckets need enough bets each. With 20 bets in a bucket, the standard error is huge and the plot is mostly noise.
  • Bucket boundaries change the picture; try a few widths or use a smoothed curve.
  • Good overall calibration can hide bad calibration in sub-groups such as favourites versus outsiders, or home versus away.
  • Testing many buckets at once invites false alarms; one bucket out of ten showing z of 2 is not surprising.
  • Calibration on training data says nothing. Only out-of-sample forecasts count.
  • Calibration alone does not make you money; the market is also well calibrated in the main Betfair football markets.

How to build it

  • Python: sklearn.calibration.calibration_curve for the reliability diagram, plus matplotlib.
  • Data: out-of-sample forecasts with results, ideally several thousand.
  • Tip: plot the market's own margin-free probabilities on the same chart; where your curve and the market's differ is where your edge or your error lives.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members