The question
"My model gives a probability for every match. How do I tell whether it's any good?"
You can't just count winners. A model that says 55% is supposed to be wrong 45% of the time. You need a score that judges the probability. The Brier score for betting models is the place to start, with log loss beside it.
The idea in one sentence
A scoring rule charges you a penalty for every forecast based on the probability you gave to what actually happened, so over many matches the model with the lower average penalty is the better forecaster.
The picture
Picture every forecast as a bet on your own honesty. If you say 70% and it comes in, you pay a small fine. If you say 70% and it doesn't, you pay a big one. Say 99% and miss, and you pay a huge one.
The two main scoring rules set those fines differently:
| You said | It happened: Brier | It didn't: Brier | It happened: log loss | It didn't: log loss |
|---|---|---|---|---|
| 50% | 0.25 | 0.25 | 0.69 | 0.69 |
| 70% | 0.09 | 0.49 | 0.36 | 1.20 |
| 90% | 0.01 | 0.81 | 0.11 | 2.30 |
| 98% | 0.0004 | 0.96 | 0.02 | 3.91 |
- The Brier score is the squared distance between your probability and the result (1 or 0). Its worst possible penalty on a yes/no question is 1.
- Log loss is minus the natural log of the probability you gave the actual result. It has no ceiling: say 2% and be wrong, and you pay 3.91, almost six times the coin-flip penalty.
Both are proper scoring rules. That means you get your best expected score by reporting what you truly believe. You can't game them by shading forecasts towards 50% or towards the extremes.
Worked Betfair example
Your Over/Under 2.5 goals model is scored over five matches against the Betfair price, with the margin already removed. (Illustrative figures. Five matches prove nothing, but they show the arithmetic.)
| Match | Model: Over | Market: Over | Result | Brier (model) | Brier (market) | Log loss (model) | Log loss (market) |
|---|---|---|---|---|---|---|---|
| 1 | 62% | 55% | Over | 0.1444 | 0.2025 | 0.478 | 0.598 |
| 2 | 48% | 52% | Under | 0.2304 | 0.2704 | 0.654 | 0.734 |
| 3 | 70% | 60% | Under | 0.4900 | 0.3600 | 1.204 | 0.916 |
| 4 | 55% | 58% | Over | 0.2025 | 0.1764 | 0.598 | 0.545 |
| 5 | 40% | 47% | Under | 0.1600 | 0.2209 | 0.511 | 0.635 |
| Average | 0.2455 | 0.2460 | 0.689 | 0.686 |
- Score match 1. Over came in, so the outcome is 1. Model Brier: (0.62 − 1)² = 0.1444. Model log loss: −ln(0.62) = 0.478.
- Score match 3. Under came in, so the outcome for "Over" is 0. Model Brier: (0.70 − 0)² = 0.49. Log loss uses the probability given to what happened, which is Under at 30%: −ln(0.30) = 1.204.
- Score the market the same way. Same formula, the market's probabilities. You now have a fair head-to-head.
- Average each column. Brier: model 0.2455, market 0.2460, so the model edges it. Log loss: model 0.689, market 0.686, so the market edges it.
- Why do they disagree? Match 3. The model was confident at 70% and wrong, and log loss punishes that harder than Brier does. That one forecast wipes out the model's lead.
- Turn it into a skill score. Brier skill against the market is 1 − 0.2455 ÷ 0.2460 = +0.2%. Positive means better than the market, but 0.2% over five matches is nothing.
Verdict: on this tiny sample the model and the market are level. To say your model beats the market, you need hundreds of matches and a clear, consistent gap on both scores.
Where does the market's probability come from?
Take the Betfair back prices, turn each into 1 ÷ odds, and scale them so they add to 100% (Lesson 2.2). For Match Odds at 1.62, 4.2 and 6.4, the raw figures add to 101.2%, and the fair probabilities are 61.0%, 23.5% and 15.5%. Use the closing price for this comparison: it's the market's best estimate.
The formula
Brier score (yes/no markets like Over/Under)
- f_i is your probability for "yes" in match i (e.g. Over 2.5).
- o_i is 1 if it happened, 0 if not.
- N is the number of matches.
In plain English: square the gap between what you said and what happened, then average. Lower is better, and saying 50% every time scores 0.25.
Brier score for Match Odds (three outcomes)
- f_j is your probability for home, draw or away.
- o_j is 1 for the result that happened and 0 for the other two.
In plain English: add up the squared miss on all three outcomes. If you said 60% home, 24% draw, 16% away and it was a draw, the score is 0.60² + 0.76² + 0.16² = 0.963. The market above (61.0/23.5/15.5) scores 0.981 on the same match.
Log loss
- p_i is the probability you gave to the result that actually happened in match i.
- ln is the natural logarithm.
In plain English: it measures how surprised your model was by what happened. Always saying 50% scores 0.693. A forecast of near-zero on something that happens is catastrophic, which is exactly the mistake that ruins a bettor.
Skill score against the market
- BS_model and BS_market are the average Brier scores on the same matches.
In plain English: above zero, you forecast better than the price; below zero, the market knows more than you. The same works with log loss in place of Brier.
Try it
Pen and paper: you rate Over 2.5 at 80% and the match ends 1-0. Work out your Brier score and your log loss, then compare them with a forecast of 50%.
Answer: Brier (0.80 − 0)² = 0.64 against 0.25; log loss −ln(0.20) = 1.61 against 0.69. Being confident and wrong costs between two and three times as much as sitting on the fence.
Common mistakes
- Scoring against a coin flip. Beating 0.25 or 0.693 is easy. The market's price is the real opponent, so score it on the same matches.
- Scoring on the matches you built the model on. In-sample scores flatter every model. Only matches the model never saw count (Lesson 7.6).
- Using the opening price as the market. The closing price carries the most information. If you only beat the opening price, you may just be slower than the market, not smarter.
- Reading a small gap as a win. Differences in the third decimal place need hundreds of matches. Work out the per-match difference and test it like any other result (Lesson 1.1).
- Trusting a good score as proof of profit. A better score means better probabilities, and that's necessary for an edge. Profit still depends on betting only where your probability beats the price by more than the 2% commission (Lesson 2.3).
Accurate models still lose when they don't beat the price: why good models still lose money.