In one sentence
Gradient boosting adds small decision trees one at a time, each nudging the prediction to fix what the previous trees got wrong.
How it works
Start with a crude guess, such as "every home side has a 50% chance of winning". Look at where that guess was worst, fit a small tree to those errors, and add a fraction of its correction. Repeat a few hundred times.
Each tree only has to explain what is left over, so together they can capture complicated patterns. The learning rate controls how big each nudge is: small steps mean more trees but less chance of lurching into noise.
This is the workhorse behind many winning tabular data competitions, via libraries like XGBoost and LightGBM. It is also the method most likely to produce a beautiful backtest that falls apart live.
The maths
- F with subscript m is the model's score (in log-odds) after m trees.
- η (eta) is the learning rate, often 0.01 to 0.1.
- h with subscript m is the new tree, fitted to the current errors.
- M is the total number of trees; p is the final probability after converting log-odds.
In words: every tree adds a small correction to a running log-odds score, and the final score is turned back into a probability.
Worked betting example
A Match Odds home-win model starts from log-odds 0, which is 50%. With learning rate 0.1, the first three trees output corrections of 1.2, 0.9 and 0.7 (illustrative) for a home side with a big rating and shot-stats advantage.
- Score after three trees = 0.1 × (1.2 + 0.9 + 0.7) = 0.28.
- Probability = 1 ÷ (1 + e to the power −0.28) ≈ 57.0%.
- After the full 300 trees the model settles on 62%.
Betfair offers 1.70 to back the home win, implying 1 ÷ 1.70 ≈ 58.8%.
- A £10 back wins £7. EV before commission = 0.62 × £7 − 0.38 × £10 = £4.34 − £3.80 = +£0.54.
- With 2% commission the win is £6.86, so EV = 0.62 × £6.86 − £3.80 ≈ +£0.45.
- Break-even after commission is 1 ÷ (1 + 0.7 × 0.98) ≈ 59.3%.
The model's edge is under three percentage points. A tiny amount of overfitting or miscalibration wipes that out, which is why the validation section below matters more than the model.
Where it's good
- Large tabular datasets: player stats, team ratings, market features, in-play state.
- Finding interactions (Elo gap × days of rest, league × home advantage) without you specifying them.
- Handles missing values natively in LightGBM and XGBoost.
- Often the most accurate single model on structured data, when validated properly.
Limitations and pitfalls
- It overfits eagerly. Training log loss keeps falling with every tree; validation log loss bottoms out and then rises. Use early stopping on a later time period, never on shuffled rows.
- Hyperparameter search (depth, learning rate, leaves, subsampling) is multiple testing in disguise. Tune enough knobs and the validation set gets overfitted too. Keep a final untouched season.
- Leakage is amplified: boosting will find any feature that secretly encodes the result, such as in-play stats joined with the wrong timestamp.
- Probabilities can be overconfident at the extremes. Check calibration before staking.
- On small datasets (a few thousand matches) a regularised logistic regression or Elo-style rating often matches or beats it out of sample.
- If the market price is one of your features, the model mostly learns the market. Compare against the market alone to see what you have added.
How to build it
- Python: lightgbm, xgboost or catboost; R: lightgbm or xgboost. Use log loss as the objective for probabilities.
- Data: several seasons, time-stamped features, results, and available odds at bet time.
- Tip: set a low learning rate (0.02-0.05), shallow trees (depth 3-5), and early stopping on the most recent season held out by date.
Related methods
- Decision trees are the pieces being added up.
- Random forests average trees in parallel and are harder to overfit.
- Regularisation ideas (penalties, shrinkage) are built into boosting's settings.
- Hyperparameter optimisation covers tuning without fooling yourself.
- Walk-forward validation is the only credible test.