Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Module 5 · Lesson 5.2

Shrinkage

“Why does 3 wins from 3 mean almost nothing?”

Intermediate9 min readBefore this: 1.2 Standard deviation and variance, 3.3 Bayes' theorem

The question

"They've gone over 2.5 goals in five of their six games. Surely the over is a bet at 1.60?"

This is where most models, and most punters, lose money quietly. Shrinkage in sports betting is the fix: a principled way to decide how much a small sample should move you away from the average.

The idea in one sentence

A team's early record is part skill, part luck, so your best estimate sits between its raw record and the league average, closer to the average when the sample is small.

The picture

Flip a fair coin three times and you'll get three heads one time in eight. In a 20-team league, that means a couple of teams will start any stat on a "perfect" run by chance alone. Most of those runs fade, not because a streak is "due" to end, but because it was never all skill in the first place.

Shrinkage is a tug-of-war between two numbers:

  • The team's own record, which is specific but noisy.
  • The league average, which is stable but not specific.

The rope is pulled by sample size. The more games you have, the more weight the team's own record earns.

Here's a team whose over 2.5 rate stays at about 79% as the season goes on, with the league average at 52% and a prior strength of k = 24 games:

Games Overs Raw rate Weight on team's record Shrunk estimate
3 3 100% 11% 57.3%
6 5 83.3% 20% 58.3%
19 15 78.9% 44% 63.9%
38 30 78.9% 61% 68.5%
76 60 78.9% 76% 72.5%

Even after a full season at 79%, the honest estimate is 68.5%. That isn't pessimism. It's what the numbers say once you remember how much a 38-game record can swing by luck.

Try it · Shrinkage demo
How strong is the pull?
League 52.0%Shrunk 58.3%Raw 83.3%0%100%
Raw rate
83.3%
Shrunk estimate
58.3%
Weight on record
20.0%
k
24.0
60%80%6 games: 58.3%120406080100
Games played, if the team keeps a 83% rate╌ raw ┈ league ━ shrunk
UsingFair oddsEV per £1 at 1.60
Raw rate1.200+£0.3233
Shrunk1.716−£0.0746
The raw rate says this is a bet; the shrunk estimate says it isn't. With only 6 games, trust the shrunk 58.3%.
The team's record gets 20% of the say. It earns more with every game.

Worked Betfair example

Over/Under 2.5 Goals. A team has gone over in 5 of its first 6 matches. The league average is 52%. Betfair offers 1.60 on over 2.5 in their next game. (Illustrative figures, and simplified: a real price also needs the opponent.)

  1. The naive estimate. 5 ÷ 6 = 83.3%. Fair odds 1 ÷ 0.833 = 1.20. At 1.60 that looks like a huge edge: 0.833 × 0.60 × 0.98 − 0.167 ≈ +£0.32 per £1.
  2. How different are teams really? Suppose past seasons show teams' true over 2.5 rates spread about τ = 10 points either side of 52%. (How to measure this is in step 7.)
  3. Prior strength. k = 0.52 × 0.48 ÷ 0.10² − 1 = 0.2496 ÷ 0.01 − 1 ≈ 24. Think of it as "the average is worth 24 games of evidence".
  4. Weight on the team's record. 6 ÷ (6 + 24) = 20%. The league average gets the other 80%.
  5. Shrunk estimate. (5 + 24 × 0.52) ÷ (6 + 24) = 17.48 ÷ 30 ≈ 58.3%. Fair odds 1 ÷ 0.583 ≈ 1.72.
  6. Expected value at 1.60 after 2% commission. 0.583 × 0.60 × 0.98 − 0.417 ≈ −£0.075 per £1. The "huge edge" was a loss of about 7.5%.
  7. Where τ comes from. Across last season, teams' over 2.5 rates over 38 games had a standard deviation of about 12.9 points. Pure luck alone would produce √(0.52 × 0.48 ÷ 38) ≈ 8.1 points of that. The true spread is what's left: √(0.129² − 0.081²) ≈ 10 points.

Verdict: at 1.60 this is a bet to leave alone. It would only become interesting at about 1.75 or bigger, where even the shrunk estimate shows a small positive EV.

The formula

The shrunk estimate

p^=x+k mn+k\hat{p} = \frac{x + k\,m}{n + k}
  • x is the number of successes (overs, wins, clean sheets).
  • n is the number of games.
  • m is the league (or group) average rate.
  • k is the prior strength, measured in games.

In plain English: pretend you've already seen k games of a perfectly average team, add the real games on top, and take the rate.

The weight on the team's own record

w=nn+k,p^=w⋅xn+(1−w) mw = \frac{n}{n + k}, \qquad \hat{p} = w \cdot \frac{x}{n} + (1 - w)\, m
  • w is the share of the estimate that comes from the team's own record.

In plain English: the same thing written as a blend. At n = k, the team's record and the average count equally.

Choosing k from the data

k=m(1−m)τ2−1,τ2=s2−m(1−m)nk = \frac{m(1-m)}{\tau^2} - 1, \qquad \tau^2 = s^2 - \frac{m(1-m)}{n}
  • τ (tau) is the spread of teams' true rates.
  • s is the spread you actually observe across teams over n games each.
  • m(1 − m) ÷ n is how much of that spread luck alone produces.

In plain English: measure how much teams differ, subtract the part luck explains, and what's left tells you how much to trust individual records. This is empirical Bayes: letting the league set its own prior.

The same idea sits under Elo's K (Lesson 5.1), under Bayes' theorem and under the penalties in regularised regression. It is also why ratings for promoted teams start near the average.

Try it

Set 5 successes from 6 games, average 52%, spread 10 points, and check the shrunk rate reads 58.3%. Then drag the games up to 38 at the same 79% rate and watch how slowly the estimate earns its way towards the raw number.

Common mistakes

  • Pricing straight from a raw rate. 5 from 6 isn't 83%. Every small-sample rate is flattered by luck, and the market knows it.
  • Shrinking to the wrong average. Shrink an over 2.5 rate toward this league's average, not all football. Shrink a newly promoted side toward promoted sides.
  • Picking k to get the answer you want. k comes from how much teams really differ, measured on past data, not from a feeling.
  • Thinking shrinkage says the streak will end. It doesn't predict a reversal. It says the evidence isn't strong enough yet, and adds weight as the evidence grows.
  • Forgetting it applies to you. Your own 20-bet hot run should be shrunk toward "no edge" just as hard. See Lesson 1.5.

Why small samples keep fooling good modellers: why good models still lose money.

Check yourself

1. League average over 2.5 goals is 52% and k = 24. A team has gone over in 3 of its first 3 games. What's the shrunk estimate?
2. What makes k (the prior strength) bigger?
3. Your model says a team has an 83% chance of over 2.5, based on 5 of 6 games. Betfair offers 1.60. What's the honest reaction?
Key takeaway

Never price from a raw small-sample rate. Blend it with the average, weighting the team's own record by n ÷ (n + k), and let the evidence earn its way in.

Go deeper in the Model Library
Next lesson
5.3 Regression: linear and logistic →
How do I turn stats into a win probability?
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members