Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Ratings and regression

TrueSkill

A Bayesian rating system that tracks each player's skill as a bell curve and updates it after each result, including team and multi-player events.

Advancedpre-matchevaluation

In one sentence

TrueSkill models every player's skill as a normal distribution with a mean and a spread, and updates both after each result using Bayesian reasoning.

How it works

Microsoft built TrueSkill to match players in online games. Each player has a mean skill, μ, and an uncertainty, σ. Think of it as "we believe this player is about μ, give or take σ".

On any given day a player performs at their skill plus random noise, called β. Whoever performs better wins. Given the two skill curves and the noise, you can work out the chance each side performs better.

After the result, the winner's mean moves up and the loser's moves down, by more when the result was a surprise and when σ was large. Both σ values shrink because you have learned something. It extends naturally to doubles, team sports and fields of more than two.

The maths

P(A beats B)=Φ ⁣(μA−μBc),c=2β2+σA2+σB2P(A \text{ beats } B) = \Phi\!\left(\frac{\mu_A - \mu_B}{c}\right), \qquad c = \sqrt{2\beta^2 + \sigma_A^2 + \sigma_B^2} μA′=μA+σA2c v(t),σA′2=σA2(1−σA2c2 w(t))\mu_A' = \mu_A + \frac{\sigma_A^2}{c}\, v(t), \qquad \sigma_A'^2 = \sigma_A^2\left(1 - \frac{\sigma_A^2}{c^2}\, w(t)\right) t=μA−μBc,v(t)=ϕ(t)Φ(t),w(t)=v(t) (v(t)+t)t = \frac{\mu_A - \mu_B}{c}, \qquad v(t) = \frac{\phi(t)}{\Phi(t)}, \qquad w(t) = v(t)\,(v(t) + t)
  • μ: a player's mean skill; σ: how uncertain we are about it.
  • β: day-to-day performance noise (default 25 ÷ 6 ≈ 4.17).
  • c: the total spread combining noise and both uncertainties.
  • Φ: the normal cumulative probability; φ: the normal density.
  • v and w: correction terms that size the update to the surprise.

In words: the win chance is the skill gap divided by the total uncertainty, and each result shifts beliefs in proportion to how surprising it was. This version ignores draws.

Worked betting example

A league match with illustrative ratings. Team A, newly promoted: μ = 28, σ = 4 (few games at this level). Team B: μ = 25, σ = 2. Use β ≈ 4.17 and leave out home advantage for simplicity.

  1. c = √(2 × 4.17² + 4² + 2²) ≈ 7.40.
  2. t = (28 − 25) ÷ 7.40 ≈ 0.41, so P(A performs better) = Φ(0.41) ≈ 65.7%.
  3. This version has no draw, so treat 65.7% as an expected score and take off half the draw rate. With 26% draws (illustrative), P(A wins) ≈ 0.657 − 0.13 ≈ 52.7%, a fair Match Odds price of about 1.90.
  4. Betfair offers 2.04. £10 stake, 2% commission: a win pays £10.40 × 0.98 = £10.19. EV ≈ 0.527 × £10.19 − 0.473 × £10 ≈ +£0.64. At 1.94 the EV would be about +£0.12, which is no bet once you allow for model error.
  5. A wins. A's μ rises to about 29.21 and σ falls to about 3.67. B's μ drops to about 24.70, and σ to about 1.96. A, the less certain player, moves much more.

Where it's good

  • Football squads rated player by player, doubles tennis, and anything where you rate individuals but results come from groups.
  • Multi-runner events such as golf or racing finishing orders, where TrueSkill can learn from full rankings.
  • Esports and darts, where results are plentiful and skill tracking matters.
  • When you want an uncertainty measure to inform stake size.

Limitations and pitfalls

  • It assumes performances are normally distributed; heavy-tailed upsets are underplayed.
  • Default parameters come from video games; β and the dynamic noise term must be retuned for each sport.
  • Team versions assume team skill is the sum of player skills, which ignores chemistry, tactics and roles.
  • It is more complex than Glicko for head-to-head sports and often not more accurate there.
  • Draw handling needs an extra draw-margin parameter that is awkward to tune for football.
  • Microsoft's original TrueSkill is patented and trademarked, so check terms before commercial use; open alternatives such as openskill exist.
  • As ever, a good rating is not an edge until it beats the closing price.

How to build it

  • Python: the trueskill package, or openskill for a permissively licensed alternative.
  • Data: results in order, with player lists for team events and finishing orders for fields.
  • Tune β and the dynamic noise τ with walk-forward log loss.
  • Tip: map conservative ratings, μ − 3σ, for leaderboards, but price matches on μ and σ together.
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members