Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Evaluation, testing and behaviour

Brier Score

A simple average squared error that scores how close your probability forecasts were to what actually happened. Lower is better.

Beginnerpre-matchin-playevaluation

In one sentence

The Brier score measures how good your probability forecasts are by averaging the squared gap between each forecast and the result, where the result is 1 if it happened and 0 if it did not.

How it works

Every time your model says "Arsenal win 60%", the match settles as either a 1 (they won) or a 0 (they did not). The Brier score takes the gap between 0.60 and that 1 or 0, squares it, and averages across all your forecasts. A perfect forecaster scores 0; someone who always says 50% on a coin-flip scores 0.25.

Squaring stops positive and negative errors cancelling out and punishes big misses more than small ones.

The score on its own means little. What matters is comparing it with a benchmark, and in betting the obvious benchmark is the market's own implied probabilities with the margin removed. If you cannot beat the market's Brier score over a decent sample, your model is unlikely to be finding value.

The maths

BS=1N∑i=1N(pi−yi)2\text{BS} = \frac{1}{N}\sum_{i=1}^{N} (p_i - y_i)^2
  • BS: the Brier score, between 0 (perfect) and 1 (every forecast fully confident and wrong).
  • N: the number of forecasts.
  • p: your forecast probability for event i, between 0 and 1.
  • y: the outcome of event i, 1 if it happened and 0 if not.

In plain English: square how far off each forecast was, then take the average.

Worked betting example

You forecast five "home win" events and compare against the market's margin-free probabilities.

Match Your p Market p Result (y) Your error² Market error²
1 0.60 0.55 1 0.1600 0.2025
2 0.30 0.35 0 0.0900 0.1225
3 0.75 0.70 1 0.0625 0.0900
4 0.20 0.25 1 0.6400 0.5625
5 0.50 0.50 0 0.2500 0.2500

Step by step for match 1: 0.60 minus 1 is −0.40, squared is 0.16.

  • Your Brier score: 1.2025 ÷ 5 = 0.2405.
  • Market Brier score: 1.2275 ÷ 5 = 0.2455.

You edge the market by 0.005 over five matches. That is far too small a sample to mean anything; the same comparison needs hundreds or thousands of events before the difference is more than noise. Note that match 4 alone, where you gave the eventual winner only 20%, contributed more than half of your total error.

Where it's good

  • Comparing two versions of your own model on the same set of matches.
  • Benchmarking a model against margin-free market prices (Betfair Starting Price or closing prices are common choices).
  • Scoring yes/no markets such as match odds for one selection, over/under 2.5 goals, or both teams to score.
  • Decomposing into calibration and sharpness, which tells you whether errors come from bias or from lack of discrimination.

Limitations and pitfalls

  • It rewards accuracy, not profit. A model can have a better Brier score than the market and still lose money after commission if its edge sits in the wrong selections.
  • Differences between good models are tiny, often in the third decimal place, so you need large samples and a significance test before drawing conclusions.
  • It is gentle on confident mistakes compared with log loss: a 99% forecast that loses costs 0.98, not an unbounded penalty. For betting, where overconfidence drives overstaking, that can hide a dangerous flaw.
  • Scores are not comparable across sports or markets, because base rates differ.
  • For outcomes with more than two ordered results (home, draw, away), a plain Brier score ignores that a draw is "closer" to a home win than an away win is; see the ranked probability score.
  • Evaluating on the same data you trained on will flatter the score badly.

How to build it

  • Python: sklearn.metrics.brier_score_loss for binary outcomes, or one line of numpy.
  • Data: your forecasts, the settled results, and a benchmark such as margin-free closing prices for the same events.
  • Tip: always report your Brier score next to the market's on the identical set of events, and bootstrap the difference to see if it is stable.
  • Log loss - a stricter score that punishes confident errors much harder.
  • Calibration - checks whether your 30% forecasts really win about 30% of the time.
  • Ranked probability score - the Brier idea extended to ordered outcomes like 1X2.
  • Model vs market - how to compare your forecasts with the exchange price fairly.
  • Implied probability - turning odds into the benchmark probabilities.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members