Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Ratings and regression

Regularisation (ridge, lasso, elastic net)

Adding a penalty for large coefficients so a model stops fitting noise, trading a little bias for much more reliable predictions.

Intermediatepre-matchin-playevaluation

In one sentence

Regularisation shrinks a model's weights towards zero by penalising their size, which reduces overfitting and usually improves predictions on new matches.

How it works

Give a regression 50 features and a few seasons of data and it will find patterns in pure noise. It will say referee X adds 4% to home wins because of three odd results. Those weights look like edges but vanish on new data.

Regularisation adds a cost for large weights. The model now has to balance fitting past results against keeping weights modest. Only features that earn their place keep sizeable weights.

Ridge shrinks every weight smoothly but rarely to exactly zero. Lasso can push weak weights to exactly zero, dropping features entirely, and elastic net mixes the two. The penalty strength, λ, is chosen by testing on data the model has not seen.

The maths

Ridge:min⁡β Loss(β)+λ∑jβj2\text{Ridge:} \quad \min_{\beta} \ \text{Loss}(\beta) + \lambda \sum_{j} \beta_j^2 Lasso:min⁡β Loss(β)+λ∑j∣βj∣\text{Lasso:} \quad \min_{\beta} \ \text{Loss}(\beta) + \lambda \sum_{j} |\beta_j| Simple case (uncorrelated, standardised features):βjridge=βj1+λ,βjlasso=sign⁡(βj)max⁡(∣βj∣−λ,0)\text{Simple case (uncorrelated, standardised features):} \quad \beta_j^{\text{ridge}} = \frac{\beta_j}{1 + \lambda}, \qquad \beta_j^{\text{lasso}} = \operatorname{sign}(\beta_j)\max\big(|\beta_j| - \lambda, 0\big)
  • Loss: how badly the model fits past data (squared error or log loss).
  • β_j: the weight on feature j (unpenalised version on the right-hand side of the simple case).
  • λ: penalty strength; zero means no regularisation.

In words: pay a price for every unit of weight, so the model only keeps big weights where the data strongly supports them. The simple-case formulas assume a particular scaling of the loss and uncorrelated features; real fits are solved numerically.

Worked betting example

A home-win logistic model with intercept 0.10 (not penalised) and three standardised features. Tonight's values: form 0.8, xG difference 0.6, referee home-bias 1.5.

  1. Unpenalised weights: 0.30, 0.45, 0.08. Log-odds = 0.10 + 0.24 + 0.27 + 0.12 = 0.73. P(home) ≈ 67.5%, fair price ≈ 1.48.
  2. Lasso with λ = 0.10: weights become 0.20, 0.35 and 0 (the referee feature is dropped). Log-odds = 0.47. P(home) ≈ 61.5%, fair price ≈ 1.63.
  3. Ridge with λ = 1.0: weights halve to 0.15, 0.225, 0.04. Log-odds = 0.415. P(home) ≈ 60.2%, fair price ≈ 1.66.
  4. Betfair offers 1.80. £10 stake at 2% commission: a win pays £8 × 0.98 = £7.84.
  5. EV: unpenalised ≈ +£2.04, lasso ≈ +£0.97, ridge ≈ +£0.74.

All three still say bet, but the overfitted model claims two to three times the edge. Stake by the unpenalised figure and Kelly-style staking will overbet badly.

Where it's good

  • Models with many features relative to the number of matches, such as player-level or referee features.
  • Team rating models with dozens of team parameters and limited games per season.
  • Machine-learning pipelines where you test many candidate features.
  • Stabilising logistic and Poisson regressions with correlated inputs.
  • Lasso as a first pass to see which features survive.

Limitations and pitfalls

  • Features must be standardised first, or the penalty hits them unfairly according to their units.
  • Choosing λ on the same data you report results on overstates performance; use walk-forward validation.
  • Lasso picks one feature from a correlated group somewhat arbitrarily; elastic net is steadier.
  • Shrunk weights are biased towards zero, so do not read them as true effect sizes.
  • Regularisation does not fix data leakage, look-ahead bias or a flawed target.
  • It will not create an edge: it only stops a model inventing one that is not there.

How to build it

  • Python: scikit-learn Ridge, Lasso, ElasticNet and LogisticRegression (penalty options), with LogisticRegressionCV for tuning; R: glmnet.
  • Data: the same as the unpenalised model, with features standardised using training data only.
  • Tune λ with time-ordered cross-validation, never random shuffles of matches.
  • Tip: regularised probabilities are usually better calibrated, which matters directly for EV and staking.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members