In one sentence
A model is well calibrated when its probabilities match reality on average: of all the selections it rates at 40%, roughly 40% win.
How it works
Group your past forecasts into buckets, say 0-10%, 10-20% and so on, then compare the average forecast in each bucket with how often those selections actually won. Plot one against the other; a perfectly calibrated model sits on the diagonal line. This plot is called a reliability diagram.
Calibration matters more for betting than almost any other property. You compare your probability with the odds to decide whether a bet has value, so if your 40% is really 31%, you will see value that is not there and bet into losses.
Calibration is not the same as being useful. A model that gives every match the league averages, say 45% home, 27% draw and 28% away, is well calibrated and tells you nothing. You want calibration and sharpness, meaning forecasts that are confident when they should be.
The maths
- ô: the observed win rate in bucket k.
- n: the number of forecasts in bucket k.
- p̄: the average forecast probability in bucket k.
- z: how many standard errors the observed rate sits from the forecast; beyond about ±2 is a warning sign.
In plain English: count how often each bucket actually won, and check whether the gap from the forecast is bigger than luck would explain.
Worked betting example
Your Match Odds model has produced 470 home-win forecasts across three buckets (illustrative figures).
| Bucket (avg forecast) | Bets | Winners | Observed rate | Standard error | z |
|---|---|---|---|---|---|
| 20% | 150 | 33 | 22.0% | 3.27% | +0.61 |
| 40% | 200 | 62 | 31.0% | 3.46% | −2.60 |
| 60% | 120 | 70 | 58.3% | 4.47% | −0.37 |
Step by step for the 40% bucket: 62 ÷ 200 = 0.31. Standard error = √(0.4 × 0.6 ÷ 200) ≈ 0.0346. The gap is −0.09, which is −2.60 standard errors, unlikely to be bad luck.
Why it matters: say a team in that bucket is available at 2.8. Your model sees expected value of 0.40 × 2.8 − 1 = +12%. If the true rate is 31%, the real expected value is 0.31 × 2.8 − 1 = −13.2%, before commission. The overconfidence has turned an apparent value bet into a clear loser.
The 20% and 60% buckets are within normal noise, so the problem is local to mid-priced teams, and a recalibration step is needed.
Where it's good
- A basic health check on any model before its probabilities go anywhere near a staking plan.
- Spotting systematic bias, such as overrating favourites or underrating draws.
- Checking the market itself: the favourite-longshot bias is a calibration failure of market prices.
- Monitoring a live model over time to catch drift after rule changes or new seasons.
Limitations and pitfalls
- Buckets need enough bets each. With 20 bets in a bucket, the standard error is huge and the plot is mostly noise.
- Bucket boundaries change the picture; try a few widths or use a smoothed curve.
- Good overall calibration can hide bad calibration in sub-groups such as favourites versus outsiders, or home versus away.
- Testing many buckets at once invites false alarms; one bucket out of ten showing z of 2 is not surprising.
- Calibration on training data says nothing. Only out-of-sample forecasts count.
- Calibration alone does not make you money; the market is also well calibrated in the main Betfair football markets.
How to build it
- Python:
sklearn.calibration.calibration_curvefor the reliability diagram, plus matplotlib. - Data: out-of-sample forecasts with results, ideally several thousand.
- Tip: plot the market's own margin-free probabilities on the same chart; where your curve and the market's differ is where your edge or your error lives.
Related methods
- Platt and isotonic scaling - the standard fixes for a miscalibrated model.
- Brier score - its decomposition includes a calibration term.
- Log loss - heavily penalises the overconfidence a calibration plot reveals.
- Favourite-longshot bias - a well-known calibration flaw in betting markets.
- Model vs market - comparing your calibration with the price's.