Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Module 5 · Lesson 5.4

Conditional logit for racing

“How do the pro racing models work?”

The question

"The big racing syndicates made fortunes from models. What were they actually doing?"

At the heart of most of them is one method: the conditional logit horse racing model. Statometrics focuses on football, but this lesson is worth your time because the method transfers directly to any market with several runners, including Match Odds and Correct Score.

The idea in one sentence

Give every runner a score built from its form, turn the scores into win chances that add up to 100%, then blend those chances with the market's prices, because the market knows things your model doesn't.

The picture

Picture five horses standing on a set of scales, each weighed by its score. A higher score means a heavier horse. Each horse's win chance is simply its share of the total weight.

That's the whole trick. It has three big advantages over modelling each horse on its own:

  • The chances always sum to 100%. Push one horse up and the others come down automatically.
  • It's "conditional" on the race. A horse is judged against today's rivals, not against every horse that ever ran.
  • Only score gaps matter. A horse 0.1 ahead of another is always e^0.1 ≈ 1.105 times as likely to win, whoever else is in the field.
Runner Score, v e^v Model chance Betfair price Market chance (margin removed)
A 1.1 3.004 33.5% 2.70 36.8%
B 1.0 2.718 30.4% 3.90 25.5%
C 0.4 1.492 16.7% 5.20 19.1%
D 0.0 1.000 11.2% 8.60 11.5%
E −0.3 0.741 8.3% 14.0 7.1%
Total 8.955 100% 100%

The scores themselves come from a weighted sum of features: speed figures, recent form, trainer and jockey records, draw, going. For example, v = 0.8 × (speed figure, standardised) + 0.4 × (trainer form) gives runner B a score of 0.8 × 1.0 + 0.4 × 0.5 = 1.0. The weights are fitted on thousands of past races, exactly like logistic regression.

Worked Betfair example

A five-runner race on the Betfair Win market. (Illustrative scores and prices, as in the table.)

  1. Turn scores into chances. Add up e^v across the field: 3.004 + 2.718 + 1.492 + 1.000 + 0.741 = 8.955. Runner B's chance is 2.718 ÷ 8.955 = 30.4%.
  2. Read the market. Betfair's back prices imply 1 ÷ 2.70 + 1 ÷ 3.90 + … = 100.7%. Divide each by 1.007 to remove the margin. B's market chance is 25.5%.
  3. The model alone. At 3.90 after 2% commission, B's EV is 0.3035 × 2.90 × 0.98 − 0.6965 ≈ +£0.166 per £1. Runner E looks like +£0.137. That's too good to be true, and it usually is.
  4. Blend with the market. Fitted on past races, the weights come out at, say, α = 0.4 on the model and β = 0.6 on the market. For B: 0.3035^0.4 × 0.2547^0.6 = 0.621 × 0.440 ≈ 0.273. Do the same for every runner, then divide each by the total so they sum to 100%.
  5. The blended chances. A 35.5%, B 27.4%, C 18.1%, D 11.4%, E 7.6%. B's fair price is now 3.65.
  6. The honest edge. At 3.90, B's EV is 0.2738 × 2.90 × 0.98 − 0.7262 ≈ +£0.052 per £1. Runner E drops to +£0.039. A, C and D are all negative.
  7. The bet. A £10 back on B pays 10 × 2.90 × 0.98 = £28.42 if it wins. Expected profit about £0.52.

Verdict: the model's +16.6% was mostly the model not knowing what the market knows. The blended +5.2% is the figure worth acting on, and even that needs to be checked against the closing price over hundreds of races.

Why the blend is the clever part

Bill Benter's well-known Hong Kong model is usually described in exactly these terms: a strong fundamental model combined with the public odds. The market price holds information no feature set captures, such as stable confidence, late fitness news and plain collective judgement. A model that ignores it is competing with it; a model that blends with it is only hunting for the gaps.

How it transfers to football

  • Match Odds is a three-runner race. Home, draw and away each get a score and the chances sum to 100%. This is a multinomial logit, the same maths.
  • Correct Score is a many-runner race. Every scoreline is a runner. Blending your Poisson grid with the market's prices uses the same α and β step.
  • The Statometrics approach. An interpretable core (Elo and Poisson) makes the call, and agreement with other independent evidence, the market above all, decides how much to trust it.

The formula

Win chance of each runner

Pi=evi∑jevj,vi=β1xi1+β2xi2+…P_i = \frac{e^{v_i}}{\sum_{j} e^{v_j}}, \qquad v_i = \beta_1 x_{i1} + \beta_2 x_{i2} + \dots
  • P_i is runner i's win chance.
  • v_i is its score, a weighted sum of its features x.
  • The sum over j runs across every runner in this race only.

In plain English: each runner gets its share of the race total, so the chances always add up to 100%.

Blending model and market

Piblend=(Pimodel)α (Pimarket)β∑j(Pjmodel)α (Pjmarket)βP_i^{\text{blend}} = \frac{(P_i^{\text{model}})^{\alpha}\,(P_i^{\text{market}})^{\beta}}{\sum_j (P_j^{\text{model}})^{\alpha}\,(P_j^{\text{market}})^{\beta}}
  • P^model is your conditional logit chance.
  • P^market is the margin-free market chance.
  • α and β are weights fitted on past races (by a second conditional logit with log P^model and log P^market as the two features).

In plain English: a weighted geometric average of the two views, re-scaled to 100%. The fitted weights tell you honestly how much your model adds beyond the market. If α comes out near zero, it adds nothing.

The key assumption

PAPB=evA−vB\frac{P_A}{P_B} = e^{v_A - v_B}
  • v_A − v_B is the gap in scores between two runners.

In plain English: the ratio of two runners' chances doesn't depend on who else is in the race. Take runner A out as a non-runner and B's chance rises to 2.718 ÷ 5.951 ≈ 45.7%, with everyone keeping their relative standing. This is usually close enough in racing, but it isn't perfect: two front-runners in the same race do take something specific from each other.

Try it

Pen and paper: using the table's scores, what's runner D's chance if runners A and B are both withdrawn?

AnswerThe remaining total is 1.492 + 1.000 + 0.741 = 3.233. D's chance is 1.000 ÷ 3.233 ≈ 30.9%, a fair price of about 3.23.

Common mistakes

  • Trusting the model without the market. A raw model edge of 15% or more almost always means the model is missing something. Blend, then decide.
  • Fitting too many features on too few races. Hundreds of features on a few thousand races will fit noise. The pros used large samples and kept testing out of sample (Lesson 7.4).
  • Treating the favourite-longshot bias as free money. Raw models tend to overrate longshots, which the blend corrects (Lesson 2.5).
  • Blending with prices you couldn't have got. Fit α and β using the price available when you would have bet, not the closing price, or the backtest cheats.
  • Thinking it's only for racing. Any market with several outcomes that sum to 100%, including Match Odds and Correct Score, fits the same frame.

How clever models still end up losing: why good models still lose money.

Check yourself

1. Why use a conditional logit rather than a separate yes/no model for each horse?
2. Your model alone makes a horse a +16.6% value bet. After blending with the market, it's +5.2%. What's the right reading?
3. Runner A (score 1.1) and runner B (score 1.0) are in the same race. How many times more likely is A to win than B?
Key takeaway

Score every runner, turn the scores into chances that sum to 100%, then blend with the market. Your model is one voice; the price is another, and usually the louder one.

Go deeper in the Model Library
Next lesson
5.5 Machine learning →
When does ML help and when does it just memorise noise?
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
MembersConditional Logit Horse Racing Model: How the Professional Racing Models Work — Statometrics Academy