The question
"My model says Over 2.5 goals is 65% in this match. When it says 65%, does it actually happen 65% of the time?"
That is the whole of betting model calibration in one line. Before you trust any value your model finds, you need to know whether its probabilities mean what they say.
The idea in one sentence
A model is calibrated when, of all the selections it rates at around 65%, about 65% actually win, and the same holds at 20%, 40% and every other level.
The picture
Take every forecast your model has made on matches it did not see while it was being built. Sort them into buckets by the probability the model gave: 30-40%, 40-50% and so on. In each bucket, work out two numbers: the average probability the model gave, and the share that actually came in.
Now plot them. The model's forecast goes along the bottom and what actually happened goes up the side. A perfectly calibrated model puts every point on the diagonal line, because 40% forecasts come in 40% of the time.
Real models wander off the line in two typical ways:
- Overconfident. The high forecasts come in less often than claimed and the low ones more often. The points form a line flatter than the diagonal. This is by far the most common fault in home-made models, and it's what overfitting produces.
- Underconfident. The opposite: the points form a line steeper than the diagonal. The model is too timid and keeps everything near the middle.
Each bucket is itself a small sample, so each point has its own wobble. The chart draws a band of 2 standard errors around the diagonal. A point inside the band could be noise; a point outside it is worth worrying about, especially if its neighbours lean the same way.
| Bucket | n | Avg forecast | Came in | Actual rate | Gap | SE | Gap in SEs |
|---|---|---|---|---|---|---|---|
| 0–10% | 50 | 7.2% | 4 | 8.0% | +0.8 pts | 3.7 pts | +0.21 |
| 10–20% | 117 | 14.9% | 17 | 14.5% | -0.4 pts | 3.3 pts | -0.12 |
| 20–30% | 112 | 24.6% | 26 | 23.2% | -1.4 pts | 4.1 pts | -0.35 |
| 30–40% | 108 | 34.7% | 31 | 28.7% | -6.0 pts | 4.6 pts | -1.30 |
| 40–50% | 113 | 44.9% | 47 | 41.6% | -3.3 pts | 4.7 pts | -0.71 |
| 50–60% | 142 | 54.9% | 66 | 46.5% | -8.5 pts | 4.2 pts | -2.03 |
| 60–70% | 97 | 65.1% | 64 | 66.0% | +0.9 pts | 4.8 pts | +0.19 |
| 70–80% | 121 | 74.9% | 94 | 77.7% | +2.8 pts | 3.9 pts | +0.72 |
| 80–90% | 96 | 84.6% | 80 | 83.3% | -1.3 pts | 3.7 pts | -0.35 |
| 90–100% | 44 | 92.5% | 41 | 93.2% | +0.7 pts | 4.0 pts | +0.18 |
Worked Betfair example
You've built an Over/Under 2.5 goals model and recorded 1,000 forecasts on matches it never saw during building. (Illustrative figures.)
| Bucket | Forecasts | Average forecast | Came in | Actual rate | Gap | Standard error | Gap in SEs |
|---|---|---|---|---|---|---|---|
| 30-40% | 150 | 35.2% | 57 | 38.0% | +2.8 | 3.9 | +0.7 |
| 40-50% | 300 | 45.1% | 126 | 42.0% | −3.1 | 2.9 | −1.1 |
| 50-60% | 330 | 54.8% | 175 | 53.0% | −1.8 | 2.7 | −0.6 |
| 60-70% | 170 | 64.6% | 97 | 57.1% | −7.5 | 3.7 | −2.1 |
| 70-80% | 50 | 73.9% | 31 | 62.0% | −11.9 | 6.2 | −1.9 |
- Work out each actual rate. In the 60-70% bucket, 97 of 170 came in: 97 ÷ 170 = 57.1%.
- Measure the gap. The model said 64.6% on average and got 57.1%, a gap of 7.5 points.
- Check it against noise. With 170 forecasts at 64.6%, the standard error is √(0.646 × 0.354 ÷ 170) = 3.7 points. The gap is 7.5 ÷ 3.7 = 2.1 standard errors, outside the band.
- Look at the pattern. The two top buckets are both well below the line, and the bottom bucket sits slightly above it. That is the overconfident shape: the model is too sure of itself at the extremes.
- See what it does to a bet. The model rates Over 2.5 at 65% and Betfair offers 1.70. Break-even after 2% commission is 1 ÷ (1 + 0.70 × 0.98) = 59.3%, so the model shouts value.
- The model's view of that bet. At 64.6%, expected value per £1 is 0.646 × 0.70 × 0.98 − 0.354 = +8.9p.
- Reality's view. At the 57.1% this bucket actually hits, it's 0.571 × 0.70 × 0.98 − 0.429 = −3.8p per £1.
Verdict: the model's "best" bets, the ones it's most excited about, are exactly where it's most wrong. Every one of them is a loser at 1.70. That is why calibration comes before staking, not after.
One number for the whole chart
The expected calibration error (ECE) weights each bucket's gap by its size: (150 × 2.8 + 300 × 3.1 + 330 × 1.8 + 170 × 7.5 + 50 × 11.9) ÷ 1,000 ≈ 3.8 points. On its own it hides where the error is, so always look at the chart too.
The formula
The actual rate in a bucket
- h_k is the number of selections in bucket k that came in.
- n_k is the number of forecasts in bucket k.
In plain English: count the winners in the bucket and divide by the number of forecasts.
The standard error of a bucket
- f̄_k is the average forecast in bucket k.
- n_k is the number of forecasts in the bucket.
In plain English: this is how far the actual rate would wander by luck if the model were perfectly calibrated. Small buckets wander a lot, so don't panic over one point outside the line.
Expected calibration error
- N is the total number of forecasts.
- |f̄_k − ō_k| is the size of the gap in bucket k, ignoring its direction.
In plain English: the average distance between what the model said and what happened, with bigger buckets counting for more.
Fixing it
A model with a consistent shape of error can be repaired by mapping its forecasts onto the actual rates, using methods such as Platt or isotonic scaling. Fit the repair on one period and check it on another, or you have just fitted the noise again.
Try it
Set the confidence factor above 1 to make the model overconfident and watch the points tip flatter than the diagonal. Then cut the number of forecasts to 200 and see how wide the noise band gets: with too few forecasts, even a bad model can look calibrated.
Common mistakes
- Checking calibration on the data you built the model on. In-sample the model always looks well calibrated. Only forecasts on matches it never saw count (Lesson 7.6).
- Reading one bucket on its own. A single point outside the band can be luck. A consistent tilt across several buckets is the real warning.
- Too few forecasts per bucket. 30 forecasts at 60% have a standard error of about 9 points. Use fewer, wider buckets until you have thousands of forecasts.
- Thinking calibrated means profitable. A model that always says what the market says is perfectly calibrated and has no edge at all. Calibration is the entry ticket; beating the price is the edge (Lesson 2.4).
- Only checking the middle. Most forecasts sit near 50%, where errors are small. The bets you actually place are usually the extreme ones, and that's where overconfidence lives.
A model that is sure of itself and wrong is the fastest way to lose money: why good models still lose money.