Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Evaluation, testing and behaviour

Log Loss

A scoring rule that punishes confident wrong forecasts very hard, closely tied to how a Kelly bettor's bank grows or shrinks.

Intermediatepre-matchin-playevaluationstaking

In one sentence

Log loss scores probability forecasts by the negative logarithm of the probability you gave to what actually happened, so being confidently wrong is punished far more than being mildly wrong.

How it works

For each event, look only at the probability you assigned to the outcome that occurred. If you said 75% and it happened, your penalty is small; if you said 20% and it happened, your penalty is much bigger. Average the penalties and you have log loss, where lower is better.

The logarithm makes the penalty explode as your probability for the actual result approaches zero. Say 1% on a winner and you pay about 4.6; say 0% and the penalty is infinite. That is why log loss is the favourite loss function for training classification models: it forces them not to be overconfident.

There is a neat betting link. If you stake using the Kelly criterion, the long-run growth rate of your bank depends on the log of the probabilities in the same way, so log loss is closer to "what happens to a Kelly bank" than the Brier score is.

The maths

LL=−1N∑i=1N[yiln⁡pi+(1−yi)ln⁡(1−pi)]\text{LL} = -\frac{1}{N}\sum_{i=1}^{N} \Big[ y_i \ln p_i + (1 - y_i)\ln(1 - p_i) \Big]
  • LL: the log loss, from 0 upwards with no ceiling.
  • N: the number of forecasts.
  • p: your forecast probability that event i happens.
  • y: 1 if event i happened, 0 if not.
  • ln: the natural logarithm.

In plain English: take the log of the probability you gave to what actually happened, flip the sign, and average.

Worked betting example

The same five home-win forecasts used on the Brier score page, against margin-free market probabilities.

Match Your p Market p Result Your penalty Market penalty
1 0.60 0.55 1 −ln 0.60 = 0.511 0.598
2 0.30 0.35 0 −ln 0.70 = 0.357 0.431
3 0.75 0.70 1 −ln 0.75 = 0.288 0.357
4 0.20 0.25 1 −ln 0.20 = 1.609 1.386
5 0.50 0.50 0 −ln 0.50 = 0.693 0.693
  • Your log loss: 3.458 ÷ 5 ≈ 0.692.
  • Market log loss: 3.465 ÷ 5 ≈ 0.693.

On Brier you led the market by 0.005; on log loss the gap almost vanishes. Match 4 is the reason: you gave the winner 20% against the market's 25%, and log loss punishes that more heavily than Brier does. The two scores can disagree about which model is better, which is worth knowing before you pick one.

For scale, a single 99% forecast that loses costs 4.61 in log loss but only 0.98 in Brier.

Where it's good

  • Training and comparing models that feed a staking plan, especially Kelly or fractional Kelly.
  • Catching overconfident models before they cause oversized stakes.
  • Multi-outcome markets (1X2, correct score, goal bands): extend by taking −ln of the probability given to the outcome that happened.
  • Correct Score, where many scorelines have small probabilities and getting the unlikely ones badly wrong is costly.

Limitations and pitfalls

  • One freak result against a near-certain forecast can dominate the average. Clip probabilities to a sensible range, such as 0.001 to 0.999, and look at the median penalty too.
  • The number is hard to interpret on its own; always compare against a baseline such as the market or a "base rate" model.
  • Like the Brier score, a better log loss does not guarantee profit once commission is taken off.
  • Small differences need large samples. Bootstrap the difference between your model and the market before believing it.
  • Outcome order is ignored, so for 1X2 it treats a draw and an away win as equally wrong when you backed the home side.
  • In-play, events are strongly correlated within a match, so thousands of in-play forecasts may carry the information of far fewer independent events.

How to build it

  • Python: sklearn.metrics.log_loss, or numpy with np.clip to avoid log of zero.
  • Data: forecasts, results, and margin-free market probabilities for the same events.
  • Tip: when training gradient boosting or logistic models, use log loss as the objective and Brier or calibration plots as a second check.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members