Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Machine learning and simulation

Decision Trees

A flowchart of yes/no questions learned from data that sorts runners or matches into groups with different win rates.

Beginnerpre-matchevaluation

In one sentence

A decision tree splits your past bets into smaller and smaller groups using simple yes/no questions, then uses the win rate in each final group as its probability.

How it works

Think of the questions a football bettor asks: is the away side rated stronger, has the home side kept a clean sheet lately, is either team playing its third game in eight days? A tree learns which question to ask first, and where to put the cut-off, by picking the split that best separates winners from losers in your data.

Each split creates two branches, and the process repeats on each branch until a stopping rule kicks in. The final groups are called leaves, and each leaf's historical win rate becomes the predicted probability for any new match that lands there.

The appeal is that you can read the whole model. The catch is that a tree will happily keep splitting until each leaf holds a handful of matches, at which point it has memorised history rather than learned anything.

The maths

G=1−∑kpk2G = 1 - \sum_{k} p_k^2 Gain=Gparent−nLnGL−nRnGR\text{Gain} = G_{\text{parent}} - \frac{n_L}{n} G_L - \frac{n_R}{n} G_R
  • G is the Gini impurity: how mixed a group is between winners and losers (0 means all one type).
  • p with subscript k is the share of the group in class k (for example winners and losers).
  • n is the number of matches in the parent group; n with L or R is the number sent left or right.
  • Gain is how much tidier the two child groups are than the parent.

In words: at every node the tree tries every feature and cut-off, and keeps the one that makes the child groups least mixed.

Worked betting example

A tree trained on league matches ends with a leaf defined as: away side rated stronger on Elo, home side conceded in each of its last five home games, and the away side had at least six days' rest. That leaf holds 120 past matches, 48 of which were away wins (illustrative figures).

  1. Leaf win rate = 48 ÷ 120 = 40%, so fair odds are 1 ÷ 0.40 = 2.5.
  2. Today a match in that leaf has the away win at 3.0 to back in Match Odds. A £10 back wins £20 or loses £10.
  3. Expected value before commission = 0.40 × £20 − 0.60 × £10 = £8 − £6 = +£2.00.
  4. With 2% commission on winnings the win is £19.60, so EV = 0.40 × £19.60 − £6 = +£1.84.
  5. Break-even probability at 3.0 after 2% commission = 1 ÷ (1 + 2 × 0.98) ≈ 33.8%.

Now the reality check. With only 120 matches, the standard error on 40% is about 4.5 points, so a rough 95% range is 31.2% to 48.8%. The break-even 33.8% sits inside that range, so this leaf does not prove an edge.

Where it's good

  • Exploring data: a shallow tree shows which factors matter and roughly where the thresholds are.
  • Explaining a model to yourself or subscribers, because every prediction has a readable path.
  • Handling mixed data (league, days of rest, Elo gap, home or away) without scaling or transformation.
  • As the building block for random forests and gradient boosting, where trees earn their keep.

Limitations and pitfalls

  • Single trees overfit badly. A deep tree can find a leaf with 12 away wins from 20 matches purely by chance, and it will look like a goldmine in-sample.
  • They are unstable: change a few matches in the training data and the top split, and therefore the whole tree, can change.
  • Probabilities come in chunky steps (one value per leaf), so they are poorly calibrated compared with logistic regression.
  • Data leakage is easy: if a feature like the closing price or the final league table sneaks in, the tree will latch onto it and your backtest will be fiction.
  • A searched tree is a form of multiple testing. Try enough features and depths and something will look profitable.
  • The market already prices most obvious splits. Recent form and team news are in the odds, so a leaf's raw win rate is not an edge unless it beats the price.

How to build it

  • Python: scikit-learn's DecisionTreeClassifier; R: rpart. Plot the tree to sanity-check every split.
  • Data: one row per match with pre-match features only, plus the result and the price you could have taken.
  • Tip: cap depth at 3-4 and require at least 100-200 matches per leaf, then test on a later season the tree never saw.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members