In one sentence
Bradley-Terry gives each competitor a positive strength and says the chance A beats B is A's strength divided by A's plus B's.
How it works
Imagine each player holds a number of raffle tickets. When two meet, one ticket is drawn from the combined pile, and whoever owns it wins. A player with three times as many tickets as another wins three times in four.
The model, published by Bradley and Terry in 1952, estimates those strengths from past head-to-head results using maximum likelihood, meaning it picks the strengths that make the observed results most probable. Because everyone is fitted together, beating strong opponents counts for more.
It is the statistical backbone behind Elo: Elo's expected-score formula is a Bradley-Terry model on a log scale, updated one match at a time. Fitting Bradley-Terry directly uses all the data at once and lets you add extras like home advantage or surface.
The maths
- π_i: the strength of competitor i (positive number).
- θ_i: the log of π_i, a rating on a scale like Elo's.
- h: an optional home or surface advantage term.
In words: odds of winning equal the ratio of the two strengths, and on the log scale the gap between ratings is all that matters, which makes it a logistic regression.
Worked betting example
Three clubs in one league, counting only games that had a winner (illustrative figures): A beat B 3 times and lost 2; A beat C 4 times and lost once; B beat C 3 times and lost twice. Draws are set aside, because the basic model has no draw.
- Fitting by maximum likelihood (a simple iterative algorithm, see below) and scaling C to 1 gives strengths A ≈ 3.16, B ≈ 1.78, C = 1.00.
- A v C: if the game has a winner, P(A wins) = 3.16 ÷ (3.16 + 1.00) ≈ 76.0%.
- Put the draw back in. Say 25% of games in this league are draws (illustrative). P(A wins) ≈ 0.75 × 76.0% ≈ 57.0%, a fair Match Odds price of about 1.75.
- Betfair offers 1.85 on A. £10 stake, 2% commission: a win pays £8.50 × 0.98 = £8.33. EV ≈ 0.570 × £8.33 − 0.430 × £10 ≈ +£0.45.
- Note the fitted strengths use every result: B's wins over C help lift A's rating because A beats B.
Fifteen games is far too few for real staking; the point is the mechanics.
Where it's good
- Football leagues, where every team meets every other, plus tennis, snooker and darts.
- Rating players or teams across leagues and tours in one consistent scale.
- Adding covariates: surface, home advantage, rest days or form become extra terms in the log-odds.
- As a sanity check on Elo: if the two disagree sharply, investigate why.
Limitations and pitfalls
- It assumes strength is fixed over the fitting window; you need time weighting or a dynamic version for form changes.
- No draws in the basic version; football needs an extension such as Davidson's tie model.
- Transitivity is baked in: if A beats B and B beats C, A must beat C, so style matchups are invisible.
- Sparse data between groups (for example, players who never meet anyone in common) gives unstable or infinite strengths without a small penalty or prior.
- Win-loss data wastes information such as goal margins or set scores.
- A model fitted on public results is unlikely to beat the Betfair price without extra inputs.
How to build it
- Python: choix, or statsmodels logistic regression on "player i minus player j" indicator columns; R: BradleyTerry2.
- Data: a list of winners and losers with dates, plus any covariates you want.
- Hunter's MM algorithm (2004) is a simple loop: new strength equals wins divided by the sum over opponents of games played ÷ combined strength.
- Tip: add a light ridge penalty so players with few matches do not get extreme ratings.
Related methods
- Elo ratings: an online, one-match-at-a-time approximation of Bradley-Terry.
- Logistic regression: the log-odds form of the same model.
- Plackett-Luce and Harville: the extension to full finishing orders.
- Massey and Colley ratings: least-squares alternatives for the same job.