In one sentence
Bayesian model averaging (BMA) prices an event as a weighted average of several models' probabilities, with each model's weight equal to how believable it is given the historical data.
How it works
Most serious bettors end up with more than one model: a Dixon-Coles goal model, an Elo-style rating, an expected-goals model. They disagree, and the temptation is to bet whenever any one of them shows value. That is a reliable way to pick up the models' errors rather than their insights.
BMA treats "which model is right?" as another uncertain quantity. Each model starts with a prior weight (often equal), and the weights are updated by how well each model predicted past matches. A model that consistently assigns higher probability to what actually happened earns more weight.
The final price is the weighted blend. Because weights are driven by likelihood, a modest difference in fit can produce a large difference in weight, so BMA often ends up leaning heavily on one model while keeping a little of the others.
The maths
- y: the outcome you are pricing, for example a home win.
- Mₖ: model number k.
- P(y given Mₖ, data): model k's probability for the outcome.
- P(Mₖ given data): model k's posterior weight.
- P(data given Mₖ): how well model k explains the past results (its likelihood).
- P(Mₖ): the prior weight on model k.
In plain English: each model votes, and the size of its vote depends on its track record.
Worked betting example
Three models price a home win (illustrative figures); Betfair offers 2.10 (implied ≈ 47.6%).
- Model views. Dixon-Coles: 48%. Elo-logistic: 52%. xG-based: 45%.
- Track record. On the same set of past matches, their log-likelihoods are −1050.0, −1051.2 and −1053.5 (higher is better). Equal prior weights.
- Weights. Weight is proportional to e raised to each log-likelihood. Relative to the best: e⁰ = 1, e to the −1.2 ≈ 0.301, e to the −3.5 ≈ 0.030. Normalised: 0.751, 0.226, 0.023.
- Blended probability. 0.751 × 48% + 0.226 × 52% + 0.023 × 45% ≈ 48.8%, fair odds ≈ 2.05.
- Decision on a £20 back at 2.10, 2% commission. BMA: 0.488 × £22 × 0.98 minus 0.512 × £20 ≈ +£0.28, a thin edge well inside model error. If you had just trusted the most bullish model (52%): ≈ +£1.61.
The blend keeps less than a fifth of the edge the Elo model claims. Most of that apparent edge is the Elo model's error.
A caution on step 3: true BMA uses each model's marginal likelihood, which also penalises complexity. Using out-of-sample log-likelihood, as here, is a common practical stand-in; in-sample figures would favour the most flexible model.
Where it's good
- Combining several reasonable pre-match models into one price.
- Avoiding "model shopping", where you pick the model that shows value on each match.
- Deciding whether a new model adds anything: if its weight stays near zero, it does not.
- Blending a market-implied model with your own, letting data decide how much to trust each.
Limitations and pitfalls
- BMA assumes one of the models is the true one. In sport none is, and the weights can collapse onto a single model even when a mix would predict better. Stacking often beats it in practice.
- Weights are very sensitive to small likelihood differences over large samples; a few lucky matches can swing them.
- Marginal likelihoods are hard to compute for complex models and very sensitive to priors.
- Correlated models (two versions of the same goal model) get double counted unless the prior weights account for it.
- Weights should be re-estimated as the season goes on, and doing this on the same data used for evaluation inflates results.
- Blending reduces error, not market efficiency. If all your models are worse than the closing price, the average will be too.
How to build it
- ArviZ has compare() with LOO-based weights ("pseudo-BMA" and stacking) for PyMC or Stan models.
- For non-Bayesian models, compute out-of-sample log loss per model in scikit-learn or pandas and turn it into weights as above.
- Data: a common hold-out set of matches on which every model is scored.
- Practical tip: include the Betfair closing price as one of the "models"; if it takes most of the weight, your models need work.
Related methods
- Ensembles and stacking – learns blend weights directly for best prediction.
- Log loss – the scoring rule behind the likelihood weights.
- Model vs market – testing whether the blend beats the price.
- MCMC – how the component Bayesian models are fitted.
- Bayes' theorem – applied here to models rather than parameters.