Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Machine learning and simulation

Ensembles and Stacking

Combines several models' probabilities, by simple averaging or a learned blend, to get a forecast better than any one alone.

Intermediatepre-matchin-playevaluation

In one sentence

An ensemble blends the forecasts of several different models, and stacking goes a step further by training a small model to decide how much to trust each one.

How it works

Three good tipsters who use different methods will each be wrong in different ways. Averaging their views cancels some of those errors, the same reason a crowd often beats an individual.

A simple ensemble just averages probabilities, perhaps with fixed weights. Stacking fits a second-level model, usually a regularised logistic regression, that learns the best weights from past predictions.

The most useful ensemble member in betting is often the market itself. Blending your model with the exchange price tends to beat either one alone, because the market holds information your model does not.

The maths

pens=∑m=1Mwm pm,∑mwm=1p_{\text{ens}} = \sum_{m=1}^{M} w_m \, p_m, \quad \sum_m w_m = 1
  • p with subscript ens is the blended probability.
  • M is the number of models; p with subscript m is model m's probability.
  • w with subscript m is model m's weight, fixed in advance or learned by stacking.

In words: the blended forecast is a weighted average of the individual forecasts. Stacking can also blend in log-odds rather than raw probability, which often works better near the extremes.

Worked betting example

Three models rate a home win:

  1. Simple average = (48 + 52 + 55) ÷ 3 ≈ 51.7%.
  2. A stacker trained on earlier seasons gives weights 0.2, 0.3 and 0.5: 0.2 × 0.48 + 0.3 × 0.52 + 0.5 × 0.55 = 52.7%.

Betfair offers 2.02, implying 1 ÷ 2.02 ≈ 49.5%.

  1. A £10 back wins £10.20. EV = 0.527 × £10.20 − 0.473 × £10 ≈ +£0.65, or about +£0.54 after 2% commission.

Now blend 50-50 with the market's 49.5%:

  1. Blended probability ≈ 51.1%. EV after 2% commission = 0.511 × £9.996 − 0.489 × £10 ≈ +£0.22.

The market blend shrinks the edge but usually makes it more real. If your model only looks good when it ignores the market, be suspicious.

Where it's good

  • Combining truly different approaches: a goals model, a rating model and a machine-learning model.
  • Blending your model with the market price, which is often the single biggest improvement available.
  • Reducing the variance of flexible models (random forests are themselves an ensemble of trees).
  • Hedging model risk: if one model breaks after a rule change, the others soften the damage.

Limitations and pitfalls

  • The classic leakage trap: training the stacker on predictions the base models made on their own training data. Those predictions are too good, so the stacker over-trusts the most overfitted model. Use out-of-fold or strictly earlier-period predictions only.
  • Stacking adds another layer of tuning, and with only a few thousand matches the learned weights are noisy. Equal weights often do as well out of sample.
  • Averaging highly similar models gains little. Three boosting models with different seeds are not three opinions.
  • Blending in probability space can pull favourites and outsiders towards the middle; log-odds blending or recalibration may be needed.
  • If one member uses closing prices and you bet earlier, the ensemble leaks future information.
  • An ensemble of weak models is still weak. It cannot create an edge that none of its members has.

How to build it

  • Python: scikit-learn StackingClassifier or VotingClassifier; mlxtend; or a hand-rolled logistic regression on out-of-fold predictions.
  • Data: time-ordered predictions from each base model, generated without the model ever seeing that match.
  • Tip: always include the market-implied probability (after margin removal) as a candidate member and compare log loss with and without your models.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members