In one sentence
Conditional logit gives every runner in a race a strength score from its features and converts the scores into win probabilities that sum to one within that race.
How it works
In a race, exactly one horse wins, so each runner's chance depends on who else is running. Conditional logit handles this by comparing runners only within the same race.
Each horse gets a score from a weighted sum of its features: speed figures, jockey, trainer, draw, days since last run and so on. Take the exponential of each score and divide by the race total, and you have win probabilities. The same weights apply to every race; the field changes.
Bill Benter's well-known approach used this in two stages: first a fundamental model from horse features, then a blend with the market's own probabilities, because public odds contain information the model misses. The blend is usually more accurate than either alone.
The maths
- V_i: runner i's score; x values: its features; β values: weights.
- p_i: the fundamental model's win probability.
- q_i: the market's implied probability with the overround removed.
- α, β: blending weights, fitted on past races.
In words: a runner's chance is its share of the field's total strength, and the final estimate blends your view with the market's.
Worked betting example
A five-runner race. Fundamental model scores V: 1.20, 0.85, 0.40, 0.10, −0.30.
- Exponentials: 3.32, 2.34, 1.49, 1.11, 0.74. Total ≈ 9.00.
- Model probabilities: 36.9%, 26.0%, 16.6%, 12.3%, 8.2%.
- Betfair back prices: 3.0, 4.2, 5.2, 6.8, 10.0. Implied total ≈ 101.1%; normalised market probabilities: 33.0%, 23.6%, 19.0%, 14.5%, 9.9%.
- Blend with α = 0.3 and β = 0.8 (market weighted more heavily): 35.6%, 24.5%, 18.0%, 13.3%, 8.7%.
- Favourite: fair price ≈ 2.81 against 3.0 available. £10 stake at 2% commission: a win pays £20 × 0.98 = £19.60. EV ≈ 0.356 × £19.60 − 0.644 × £10 ≈ +£0.54.
- Second favourite: blended 24.5% at 4.2 gives EV ≈ 0.245 × £31.36 − 0.755 × £10 ≈ +£0.13, a thin edge. The other three runners are all negative.
Notice how the blend pulls the model's 36.9% back towards the market's 33.0%. That shrinking is deliberate.
Where it's good
- Horse and greyhound racing win markets, its classic home.
- Any event with one winner from a varying field: golf outright, to-score-first markets.
- Combining your model with market prices in a principled way.
- Feeding place and forecast pricing through Plackett-Luce style extensions.
Limitations and pitfalls
- It assumes independence of irrelevant alternatives: removing one runner scales everyone else up proportionally, which ignores pace and tactical interactions.
- Fitting needs thousands of races with clean, consistent features.
- Blending with the market means your edge is only the part the market misses, which is usually small.
- The favourite-longshot bias means raw implied probabilities misprice outsiders; the blend weights partly correct this.
- Non-runners and late market moves change q; stale prices produce phantom value.
- Benter's reported success came from a large team, huge data effort and a pool-betting market; do not expect the same result from a weekend project.
How to build it
- Python: statsmodels ConditionalLogit, pylogit, or a custom likelihood in PyTorch; R: survival::clogit or mlogit.
- Data: every runner in every race with features and the winner flagged, grouped by race ID.
- Fit the fundamental model and the blend on separate time periods to avoid leakage.
- Tip: use Betfair starting prices or late exchange prices as q, not early prices, when fitting the blend.
Related methods
- Multinomial and ordinal logistic regression: the fixed-outcome relatives.
- Plackett-Luce and Harville: extend win probabilities to full finishing orders.
- Speed figures: a key input feature.
- Model vs market: testing whether your blend adds anything.
- Favourite-longshot bias: why raw market odds need adjusting.