In one sentence
K-nearest neighbours (kNN) finds the k past matches that look most like today's and uses their results as the forecast.
How it works
Punters do this informally: "last time a mid-table side with this xG profile hosted a team at these odds, they won." kNN makes it systematic by measuring the distance between matches on the features you choose.
Pick a number k, say 5 or 50. For a new match, find the k closest past matches and take the share that ended in a home win as your probability. You can weight closer matches more heavily.
There is no training step, which makes kNN easy to understand. It also makes it fragile: the answer depends heavily on which features you use, how you scale them, and how many neighbours you pick.
The maths
- d(a, b) is the distance between matches a and b.
- a with subscript f is match a's value on feature f; s with subscript f is a scale for that feature so none dominates.
- k is the number of neighbours; y with subscript i is 1 if neighbour i was a home win, 0 if not.
- p-hat is the predicted probability.
In words: find the k most similar matches after putting features on a common scale, and average their results.
Worked betting example
Today's match: home side's xG difference over the last six games is +0.6, and the home implied probability is 50% (2.00 on Betfair). Scale xG by 0.5 and implied probability by 0.05.
Eight past matches (xG diff, implied probability, home win yes or no) and their distances:
- 0.7, 49%, yes: distance 0.28
- 0.5, 52%, yes: distance 0.45
- 0.4, 47%, no: distance 0.72
- 0.6, 45%, no: distance 1.00
- 0.9, 55%, yes: distance 1.17
- 0.1, 40%, no: distance 2.24
- 1.2, 62%, yes: distance 2.68
- 0.3, 35%, no: distance 3.06
- With k = 5 the nearest five are the first five rows: 3 home wins from 5, so p = 60%.
- A £10 back at 2.00 wins £10. EV = 0.60 × £10 − 0.40 × £10 = +£2.00, or +£1.88 after 2% commission.
Here is the problem. The standard error of a proportion from 5 matches is √(0.6 × 0.4 ÷ 5) ≈ 22 percentage points. The honest answer is "somewhere between about 17% and 100%", which tells you nothing. You need hundreds of neighbours for a usable estimate, and then they stop being very similar.
Where it's good
- A quick sanity check: "how did matches like this actually go?"
- Finding comparable matches for manual analysis.
- Low-dimensional problems with lots of data, such as in-play states defined by score, time and pre-match price.
- A nonparametric benchmark to test whether a fancier model adds anything.
Limitations and pitfalls
- Small k gives wildly noisy probabilities; large k blurs everything towards the average. Choose k by out-of-sample log loss, not by eye.
- Scaling decides the answer. Leave odds in raw units and xG in raw units and one will swamp the other.
- The curse of dimensionality: with more than a handful of features, nothing is truly near anything. Consider dimensionality reduction first.
- Leakage: make sure neighbours are only drawn from matches before the one you are predicting.
- If market odds are a feature, kNN mostly echoes the market. Check it adds value beyond the price.
- A logistic regression on the same features usually gives smoother, better-calibrated probabilities.
How to build it
- Python: scikit-learn KNeighborsClassifier with StandardScaler; R: the class or kknn package.
- Data: a few well-chosen pre-match features and results; more rows matter more than more columns.
- Tip: use distance weighting and choose k with time-ordered cross-validation; report the neighbour count behind every prediction.
Related methods
- Clustering groups similar matches without a target outcome.
- Dimensionality reduction makes distances meaningful with many features.
- Logistic regression is the smoother baseline.
- P-values and confidence intervals explain why 3 from 5 means little.
- Binomial gives the maths behind small-sample win rates.