In one sentence
A random forest grows many deliberately different decision trees and averages their answers, which cancels out much of the noise any single tree picks up.
How it works
One tipster who studies every past match will overfit to quirks. Ask 500 tipsters who each saw a random sample of matches and a random handful of stats, and their average view is far steadier.
That is a random forest. Each tree is trained on a bootstrap sample (rows drawn with replacement) and, at each split, only considers a random subset of features. The trees disagree in useful ways, and averaging smooths out their individual mistakes.
Because each tree skips about a third of the rows, you can score every row using only the trees that never saw it. This out-of-bag score is a handy free check, though it does not replace testing on later seasons.
The maths
- p-hat of x is the forest's probability for match x.
- B is the number of trees, typically 300 to 1,000.
- p-hat with subscript b is the probability from tree b, taken from the leaf the match lands in.
In words: the forest's answer is simply the average of all the trees' answers.
Worked betting example
A forest trained on five seasons of league football predicts Over 2.5 goals for a match at 58%. Betfair offers 1.80 to back Over 2.5, which implies 1 ÷ 1.80 ≈ 55.6%.
- A £20 back at 1.80 wins £16 or loses £20.
- EV before commission = 0.58 × £16 − 0.42 × £20 = £9.28 − £8.40 = +£0.88.
- With 2% commission the win is £15.68, so EV = 0.58 × £15.68 − £8.40 ≈ +£0.69.
Now check calibration. Suppose the out-of-bag results show that matches the forest rates at 58% actually went over only 54% of the time.
- At a true 54%, EV = 0.54 × £16 − 0.46 × £20 = £8.64 − £9.20 = −£0.56 before commission, and about −£0.73 after 2%.
A four-point calibration error turned a small edge into a loss. That gap is common with forests, which pull probabilities towards the middle.
Where it's good
- Tabular pre-match data with dozens of features and messy interactions (form, fixtures congestion, travel, weather).
- A strong baseline with little tuning: default settings are often close to the best a forest can do.
- Ranking features by importance to see what the data is leaning on.
- Robust to outliers and does not need features scaled.
Limitations and pitfalls
- Forests shrink extreme probabilities towards the middle, so short favourites look too long and outsiders too short. Run calibration checks and consider Platt or isotonic scaling.
- Out-of-bag scores assume rows are independent. Matches from the same week share information, so OOB results flatter the model versus a proper walk-forward test.
- Feature importance is biased towards features with many distinct values, such as IDs or dates. A team ID ranking top is a warning sign, not an insight.
- Leakage kills: including closing prices, line-ups confirmed after your bet time, or season-end tables will make any backtest look superb.
- On many betting datasets a well-built logistic regression with a few good features matches a forest. Use the forest if it wins out of sample, not because it sounds cleverer.
- Beating the closing line is the real test. A forest that beats 55.6% on paper but cannot beat the market price after commission is not an edge.
How to build it
- Python: scikit-learn RandomForestClassifier; R: ranger. Set min_samples_leaf to at least 20-50 so leaves are not tiny.
- Data: pre-match features timestamped before your bet time, results, and the odds available at that moment.
- Tip: split train and test by date, never randomly, and report log loss alongside profit.
Related methods
- Decision trees are the building block.
- Gradient boosting usually edges forests on accuracy but overfits more easily.
- Calibration tells you whether 58% really means 58%.
- Platt and isotonic scaling fix miscalibrated forest outputs.
- Walk-forward validation is the honest replacement for out-of-bag scores.