Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Evaluation, testing and behaviour

Information Theory

Measures uncertainty and the difference between your probabilities and the market's, and links that difference directly to how fast a Kelly bank can grow.

Advancedpre-matchevaluationstaking

In one sentence

Information theory gives you tools, entropy and KL divergence, to measure how uncertain an event is and how much better your probabilities are than the market's, in units that translate into bank growth.

How it works

Entropy measures uncertainty. A match between two evenly matched sides has high entropy; a title contender at home to the bottom side has low entropy. It tells you how much there is to learn about an event before it happens.

KL divergence, named after Kullback and Leibler, measures how different two sets of probabilities are, such as your model's and the market's. It is zero when they agree and grows as they disagree.

The link to betting was set out by Kelly in 1956. If your probabilities are correct and the market's odds are fair (no margin), the best possible long-run growth rate of your bank using Kelly staking equals the KL divergence between your probabilities and the market's. Edge, in this view, is information the market does not have.

The maths

H(p)=−∑ipilog⁡2piH(p) = -\sum_{i} p_i \log_2 p_i DKL(p ∥ q)=∑ipiln⁡piqiD_{KL}(p \,\|\, q) = \sum_{i} p_i \ln\frac{p_i}{q_i}
  • H: entropy, in bits when using log base 2.
  • p: your (assumed true) probability for outcome i.
  • q: the market's implied probability for outcome i, with the margin removed.
  • D KL: the KL divergence from q to p, in natural-log units.
  • ln: the natural logarithm.

In plain English: KL divergence adds up, across outcomes, how much more likely you think each outcome is than the market does, weighted by how often it happens; with fair odds this is your maximum expected log growth per event.

Worked betting example

A football match. Your model says home 50%, draw 30%, away 20%. The fair (margin-free) market prices are about 2.22, 3.33 and 4.00, implying 45%, 30% and 25%.

  1. Entropy of your forecast: about 1.49 bits, a fairly open match.
  2. KL divergence: 0.50 × ln(0.50 ÷ 0.45) + 0.30 × ln(0.30 ÷ 0.30) + 0.20 × ln(0.20 ÷ 0.25) ≈ 0.0527 + 0 − 0.0446 ≈ 0.0081.

With a £100 bank and Kelly's approach for fair odds on all outcomes, you split the bank in proportion to your probabilities: £50 on home, £30 on the draw, £20 on away.

  • Home wins: £50 × 2.222 ≈ £111.11.
  • Draw: £30 × 3.333 ≈ £100.00.
  • Away wins: £20 × 4.00 = £80.00.

Expected log growth = 0.5 × ln 1.111 + 0.3 × ln 1.000 + 0.2 × ln 0.800 ≈ 0.0081, exactly the KL divergence. That is about 0.81% growth per match in the long-run, typical sense, if your probabilities are right. In practice commission and the market margin cut into this, and any error in your probabilities reduces it or turns it negative.

Where it's good

  • Understanding why log loss is the natural score for betting models: it is the cross-entropy, and differences in it map to differences in achievable growth.
  • Measuring how far your model disagrees with the market, event by event, to spot where your edge is claimed to come from.
  • Correct Score and other many-outcome markets, where KL divergence handles every scoreline cleanly.
  • Feature selection, using mutual information to see which inputs carry information about results.

Limitations and pitfalls

  • The growth result assumes your probabilities are the truth. If they are wrong, a large KL divergence means large disagreement, not large edge, and Kelly staking will overbet.
  • Commission and the market margin reduce the achievable growth; the clean equality only holds for fair odds.
  • Proportional betting on every outcome is a teaching case; in practice you back only the outcomes where you see value, and Betfair commission applies to net winnings per market.
  • Numbers are small and abstract, so it is easy to misread them. Convert to growth per event or per hundred events.
  • Estimating entropy or mutual information from small samples is biased; use plenty of data.
  • Useful for understanding and diagnostics more than as a trading signal on its own.

How to build it

  • Python: scipy.stats.entropy computes both entropy and KL divergence; sklearn.feature_selection.mutual_info_classif for mutual information.
  • Data: your probabilities and margin-free market probabilities for the same events.
  • Tip: plot the per-event KL divergence against realised profit; if the events where you disagree most are not where you profit most, your model's confidence is misplaced.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members