Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Evaluation, testing and behaviour

Platt and Isotonic Scaling

Two ways to repair a model whose probabilities are biased, by learning a mapping from raw scores to realistic probabilities on held-out data.

Intermediatepre-matchin-playevaluation

In one sentence

Platt scaling and isotonic regression are recalibration methods: they take your model's raw probabilities and bend them so that forecasts of 40% really do win about 40% of the time.

How it works

Many models, especially tree ensembles, neural networks and support vector machines, produce scores that rank selections well but are off as probabilities. Rather than rebuild the model, you fit a small second model that translates its output into calibrated probabilities, using data the first model never saw.

Platt scaling fits an S-shaped logistic curve with just two numbers: a slope and a shift. If the slope comes out below 1, your model is overconfident and the curve pulls extreme forecasts back towards the middle. It is stable with small datasets but can only fix smooth, S-shaped distortions.

Isotonic regression fits a staircase that only ever goes up, so higher raw scores never map to lower probabilities. It can fix odd, lumpy miscalibration, but it needs a lot more data and can overfit badly with only a few hundred results.

The maths

pcal=11+e−(a⋅s+b),s=ln⁡praw1−prawp_{\text{cal}} = \frac{1}{1 + e^{-(a \cdot s + b)}}, \qquad s = \ln\frac{p_{\text{raw}}}{1 - p_{\text{raw}}}
  • p cal: the recalibrated probability.
  • p raw: the model's original probability.
  • s: the log-odds of the raw probability, which stretches probabilities onto an unbounded scale.
  • a: the slope; below 1 shrinks overconfident forecasts, above 1 sharpens timid ones.
  • b: the shift; negative values push all probabilities down a little.

In plain English: convert the forecast to log-odds, scale and shift it, and convert back. Isotonic regression has no neat formula; it is the best-fitting non-decreasing step function through the held-out results.

Worked betting example

Your football over 2.5 goals model is overconfident. On a held-out season, Platt scaling finds a slope of a = 0.8 and shift b = −0.1.

Take a match your model rates at 80% for over 2.5 goals.

  1. Log-odds: ln(0.8 ÷ 0.2) = ln 4 ≈ 1.386.
  2. Scale and shift: 0.8 × 1.386 − 0.1 ≈ 1.009.
  3. Back to probability: 1 ÷ (1 + e to the −1.009) ≈ 0.733.

So 80% becomes about 73.3%. Across the range:

Raw forecast Calibrated
20% 23.0%
40% 39.5%
60% 55.6%
80% 73.3%

Now suppose over 2.5 is trading at 1.30 on Betfair, implying 76.9%. The raw model saw value: 0.80 × 1.30 − 1 = +4% before commission. The calibrated model sees 0.733 × 1.30 − 1 ≈ −4.7%, a clear no-bet. That single correction stops a steady stream of short-priced losing bets.

Where it's good

  • Tree-based models (random forests, gradient boosting) whose probabilities cluster away from 0 and 1.
  • Models trained with class weights or resampling, which distort probabilities by design.
  • Quick fixes when a model ranks selections well but the calibration plot is off.
  • Isotonic scaling for large datasets, such as many seasons across many football leagues, where the distortion is irregular.

Limitations and pitfalls

  • The calibration data must be separate from the training data. Fitting on the same data makes the correction useless and often harmful.
  • Isotonic regression overfits with small samples, producing flat steps and sudden jumps; below a few thousand results, prefer Platt.
  • Both are one-dimensional: they cannot fix a model that is overconfident on away teams but fine on home teams. Fit separate calibrators per segment if needed, with enough data in each.
  • In three-way markets (home, draw, away), calibrating each outcome separately means the probabilities no longer sum to 1; renormalise afterwards.
  • Recalibration cannot add information. It improves probabilities, not the model's ability to separate winners from losers.
  • Calibration drifts. A curve fitted on 2022 may be wrong by 2025, so refit on a rolling window.

How to build it

  • Python: sklearn.calibration.CalibratedClassifierCV with method set to sigmoid (Platt) or isotonic; or sklearn.isotonic.IsotonicRegression directly.
  • Data: a dedicated held-out calibration set, or cross-validated predictions, with at least several hundred results for Platt.
  • Tip: use walk-forward splits so the calibrator only ever learns from earlier matches than those it is applied to.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members