Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Module 7 · Lesson 7.1

Calibration

“When I say 30%, does it happen 30% of the time?”

The question

"My model says Over 2.5 goals is 65% in this match. When it says 65%, does it actually happen 65% of the time?"

That is the whole of betting model calibration in one line. Before you trust any value your model finds, you need to know whether its probabilities mean what they say.

The idea in one sentence

A model is calibrated when, of all the selections it rates at around 65%, about 65% actually win, and the same holds at 20%, 40% and every other level.

The picture

Take every forecast your model has made on matches it did not see while it was being built. Sort them into buckets by the probability the model gave: 30-40%, 40-50% and so on. In each bucket, work out two numbers: the average probability the model gave, and the share that actually came in.

Now plot them. The model's forecast goes along the bottom and what actually happened goes up the side. A perfectly calibrated model puts every point on the diagonal line, because 40% forecasts come in 40% of the time.

Real models wander off the line in two typical ways:

  • Overconfident. The high forecasts come in less often than claimed and the low ones more often. The points form a line flatter than the diagonal. This is by far the most common fault in home-made models, and it's what overfitting produces.
  • Underconfident. The opposite: the points form a line steeper than the diagonal. The model is too timid and keeps everything near the middle.

Each bucket is itself a small sample, so each point has its own wobble. The chart draws a band of 2 standard errors around the diagonal. A point inside the band could be noise; a point outside it is worth worrying about, especially if its neighbours lean the same way.

Try it · Calibration plot
0%0%20%20%40%40%60%60%80%80%100%100%Average forecast
Up: how often it came in · dashed line: perfect calibration · faint line: the truth you set
Forecasts
1,000
Average gap (ECE)
3.1 pts
Brier score
0.183
Verdict
Some buckets out of line: look closer
1 bucket is more than 2 standard errors off the diagonal (amber dots).
BucketnAvg forecastCame inActual rateGapSEGap in SEs
0–10%507.2%48.0%+0.8 pts3.7 pts+0.21
10–20%11714.9%1714.5%-0.4 pts3.3 pts-0.12
20–30%11224.6%2623.2%-1.4 pts4.1 pts-0.35
30–40%10834.7%3128.7%-6.0 pts4.6 pts-1.30
40–50%11344.9%4741.6%-3.3 pts4.7 pts-0.71
50–60%14254.9%6646.5%-8.5 pts4.2 pts-2.03
60–70%9765.1%6466.0%+0.9 pts4.8 pts+0.19
70–80%12174.9%9477.7%+2.8 pts3.9 pts+0.72
80–90%9684.6%8083.3%-1.3 pts3.7 pts-0.35
90–100%4492.5%4193.2%+0.7 pts4.0 pts+0.18
Each dot is a bucket of forecasts: across is what the model said on average, up is how often it actually came in. Dots are sized by how many forecasts they hold. Inside the shaded band, the gap could just be luck.

Worked Betfair example

You've built an Over/Under 2.5 goals model and recorded 1,000 forecasts on matches it never saw during building. (Illustrative figures.)

Bucket Forecasts Average forecast Came in Actual rate Gap Standard error Gap in SEs
30-40% 150 35.2% 57 38.0% +2.8 3.9 +0.7
40-50% 300 45.1% 126 42.0% −3.1 2.9 −1.1
50-60% 330 54.8% 175 53.0% −1.8 2.7 −0.6
60-70% 170 64.6% 97 57.1% −7.5 3.7 −2.1
70-80% 50 73.9% 31 62.0% −11.9 6.2 −1.9
  1. Work out each actual rate. In the 60-70% bucket, 97 of 170 came in: 97 ÷ 170 = 57.1%.
  2. Measure the gap. The model said 64.6% on average and got 57.1%, a gap of 7.5 points.
  3. Check it against noise. With 170 forecasts at 64.6%, the standard error is √(0.646 × 0.354 ÷ 170) = 3.7 points. The gap is 7.5 ÷ 3.7 = 2.1 standard errors, outside the band.
  4. Look at the pattern. The two top buckets are both well below the line, and the bottom bucket sits slightly above it. That is the overconfident shape: the model is too sure of itself at the extremes.
  5. See what it does to a bet. The model rates Over 2.5 at 65% and Betfair offers 1.70. Break-even after 2% commission is 1 ÷ (1 + 0.70 × 0.98) = 59.3%, so the model shouts value.
  6. The model's view of that bet. At 64.6%, expected value per £1 is 0.646 × 0.70 × 0.98 − 0.354 = +8.9p.
  7. Reality's view. At the 57.1% this bucket actually hits, it's 0.571 × 0.70 × 0.98 − 0.429 = −3.8p per £1.

Verdict: the model's "best" bets, the ones it's most excited about, are exactly where it's most wrong. Every one of them is a loser at 1.70. That is why calibration comes before staking, not after.

One number for the whole chart

The expected calibration error (ECE) weights each bucket's gap by its size: (150 × 2.8 + 300 × 3.1 + 330 × 1.8 + 170 × 7.5 + 50 × 11.9) ÷ 1,000 ≈ 3.8 points. On its own it hides where the error is, so always look at the chart too.

The formula

The actual rate in a bucket

oˉk=hknk\bar{o}_k = \frac{h_k}{n_k}
  • h_k is the number of selections in bucket k that came in.
  • n_k is the number of forecasts in bucket k.

In plain English: count the winners in the bucket and divide by the number of forecasts.

The standard error of a bucket

SEk=fˉk(1−fˉk)nk\text{SE}_k = \sqrt{\frac{\bar{f}_k (1 - \bar{f}_k)}{n_k}}
  • f̄_k is the average forecast in bucket k.
  • n_k is the number of forecasts in the bucket.

In plain English: this is how far the actual rate would wander by luck if the model were perfectly calibrated. Small buckets wander a lot, so don't panic over one point outside the line.

Expected calibration error

ECE=∑knkN∣fˉk−oˉk∣\text{ECE} = \sum_{k} \frac{n_k}{N} \left| \bar{f}_k - \bar{o}_k \right|
  • N is the total number of forecasts.
  • |f̄_k − ō_k| is the size of the gap in bucket k, ignoring its direction.

In plain English: the average distance between what the model said and what happened, with bigger buckets counting for more.

Fixing it

A model with a consistent shape of error can be repaired by mapping its forecasts onto the actual rates, using methods such as Platt or isotonic scaling. Fit the repair on one period and check it on another, or you have just fitted the noise again.

Try it

Set the confidence factor above 1 to make the model overconfident and watch the points tip flatter than the diagonal. Then cut the number of forecasts to 200 and see how wide the noise band gets: with too few forecasts, even a bad model can look calibrated.

Common mistakes

  • Checking calibration on the data you built the model on. In-sample the model always looks well calibrated. Only forecasts on matches it never saw count (Lesson 7.6).
  • Reading one bucket on its own. A single point outside the band can be luck. A consistent tilt across several buckets is the real warning.
  • Too few forecasts per bucket. 30 forecasts at 60% have a standard error of about 9 points. Use fewer, wider buckets until you have thousands of forecasts.
  • Thinking calibrated means profitable. A model that always says what the market says is perfectly calibrated and has no edge at all. Calibration is the entry ticket; beating the price is the edge (Lesson 2.4).
  • Only checking the middle. Most forecasts sit near 50%, where errors are small. The bets you actually place are usually the extreme ones, and that's where overconfidence lives.

A model that is sure of itself and wrong is the fastest way to lose money: why good models still lose money.

Check yourself

1. Your model rated 200 selections at around 40%. 62 of them won. What does that suggest?
2. Why does calibration matter more to a bettor than simply picking more winners?
3. A bucket of 50 forecasts at 74% came in at 62%. A bucket of 330 forecasts at 55% came in at 53%. Which gap is stronger evidence of a problem?
Key takeaway

Your model's 65% is only worth betting if it happens about 65% of the time. Check it bucket by bucket, on bets the model never saw, before you trust a single value call.

Go deeper in the Model Library
Next lesson
7.2 Brier score and log loss →
How do I score a probability model?
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members