The question
"I've got xG, shots, form and Elo. How do I turn all that into one number: the chance the home side wins?"
That's exactly the job of regression. This lesson covers logistic regression for football betting, and its simpler cousin linear regression, so you can build a price from stats and understand what the numbers inside it mean.
The idea in one sentence
Regression finds the weights that best link your inputs to past outcomes: linear regression for amounts (goals, shots), logistic regression for chances (win, over 2.5, both teams to score).
The picture
Plot last season's matches on a chart. The x-axis is the xG gap: the home side's average xG difference per game minus the away side's. The y-axis is what happened: 1 if the home side won, 0 if not.
Every dot sits on either the top or the bottom line. Regression's job is to draw the curve through them that best says "at this gap, this share of home sides won".
- A straight line (linear regression) fits fine in the middle but keeps going. At a gap of +4 it predicts a 125% win chance, which is impossible.
- An S-curve (logistic regression) bends at both ends. It gets close to 0% and 100% but never crosses them.
Here's the same illustrative fit both ways (home win base rate 45%):
| xG gap | Straight line | S-curve |
|---|---|---|
| −1.0 | 25% | 25.0% |
| 0 | 45% | 45.0% |
| +0.5 | 55% | 56.2% |
| +1.0 | 65% | 66.8% |
| +2.0 | 85% | 83.2% |
| +4.0 | 125% | 96.8% |
In the middle, the two agree. At the extremes, only the S-curve makes sense. That's why chances are always modelled with the logistic form.
Worked Betfair example
Match Odds, back the home side. A logistic model has been fitted on past seasons (illustrative coefficients):
- Intercept β₀ = −0.20. With a gap of zero the home side wins 45% of the time, and the log-odds of 45% is −0.20.
- Slope β₁ = 0.90 per goal of xG gap.
- The input. The home side averages +0.4 xG difference per game; the away side −0.1. The gap is 0.4 − (−0.1) = +0.5.
- The linear score (log-odds). z = −0.20 + 0.90 × 0.5 = 0.25.
- Turn it into a chance. p = 1 ÷ (1 + e^(−0.25)) = 1 ÷ (1 + 0.779) ≈ 56.2%.
- Fair price. 1 ÷ 0.562 ≈ 1.78.
- Betfair price. The home side is 1.85 to back, which implies 54.1%.
- Expected value after 2% commission. 0.562 × 0.85 × 0.98 − 0.438 ≈ +£0.030 per £1, about +3%.
- Stake and payout. A £10 back at 1.85 pays 10 × 0.85 × 0.98 = £8.33 profit if it wins, and loses £10 if it doesn't.
Verdict: a small edge on paper. The coefficients were estimated from a finite sample, so each one carries its own uncertainty. Before trusting the +3%, you want the model to be calibrated on matches it has never seen, and you want your prices to beat the close.
Linear regression's job: predicting amounts
Linear regression is still the right tool when the answer is an amount. Say a fitted model predicts total goals as 0.35 + 0.90 × (combined xG of both sides). For a combined xG of 2.9 that gives 0.35 + 0.90 × 2.9 = 2.96 goals.
You can then feed 2.96 into a Poisson model: the chance of 3 or more goals is about 56.8%, a fair over 2.5 price of about 1.76. Linear for the amount, Poisson for the chance: two simple, interpretable steps.
The formula
Linear regression
- y is the amount you're predicting, such as total goals.
- x₁, x₂, … are the inputs, such as xG or rating gap.
- β₀ is the intercept (the prediction when every input is zero), and β₁, β₂, … are the weights.
- ε is the part the inputs can't explain: luck, mostly.
In plain English: each input adds its weight times its value, and the fit chooses the weights that make the squared misses as small as possible across past matches.
Logistic regression
- p is the chance of the event (home win, over 2.5).
- z is the linear score, also called the log-odds.
- e is the constant 2.718…
In plain English: add up the inputs just like linear regression, then squash the total through an S-curve so the answer is always a valid chance.
Reading a coefficient
- p ÷ (1 − p) is the odds in "chances for : chances against" form, not Betfair's decimal odds.
In plain English: a coefficient of 0.90 multiplies the odds of a home win by e^0.90 ≈ 2.46 for each extra goal of xG gap. It does not add a fixed number of percentage points, because the curve flattens near 0% and 100%.
Three outcomes, not two
Match Odds has three results, so a single yes/no model only prices one of them. The usual fixes are a multinomial or ordinal logistic model, which treats home, draw and away together, or a goals model that produces all three at once. For Statometrics, the core that makes the call is the interpretable Elo and Poisson pair, and regression is how you fit and check their inputs.
Keeping it honest
- Few, sensible inputs. Every extra input improves the fit to the past and risks fitting noise. Regularisation shrinks weak coefficients toward zero, the same idea as Lesson 5.2.
- Only information known before kick-off. An input built with any later data makes the backtest useless.
- Fit on the past, test on the future. Score the model on seasons it never saw (Lesson 7.6).
Try it
Pen and paper: with β₀ = −0.20 and β₁ = 0.90, what's the home win chance and fair price when the xG gap is +1.0? Is 1.55 on Betfair a bet after 2% commission?
Answer
z = −0.20 + 0.90 × 1.0 = 0.70. p = 1 ÷ (1 + e^(−0.70)) = 1 ÷ 1.497 ≈ 66.8%, a fair price of about 1.50. At 1.55 the EV is 0.668 × 0.55 × 0.98 − 0.332 ≈ +£0.028 per £1, so a small edge on paper.Common mistakes
- Using a straight line for a chance. It will produce impossible prices at the extremes, exactly where you're most tempted to bet.
- Reading coefficients as percentage points. Logistic weights act on log-odds. The same input moves a 50% chance much more than a 90% one.
- Throwing in every stat you have. Shots, corners, possession and xG overlap heavily. Correlated inputs give unstable weights that flip from season to season.
- Judging the model on its own training data. In-sample fit always looks good. Only matches it hasn't seen tell you anything (Lesson 7.4).
- Forgetting the market is an input too. A model that ignores the price is competing with everyone who's already looked at the same stats.
The bigger picture on why fitted models disappoint: why good models still lose money.