Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Machine learning and simulation

K-Nearest Neighbours

Predicts an outcome by finding the most similar past matches and seeing how they turned out.

Beginnerpre-matchevaluation

In one sentence

K-nearest neighbours (kNN) finds the k past matches that look most like today's and uses their results as the forecast.

How it works

Punters do this informally: "last time a mid-table side with this xG profile hosted a team at these odds, they won." kNN makes it systematic by measuring the distance between matches on the features you choose.

Pick a number k, say 5 or 50. For a new match, find the k closest past matches and take the share that ended in a home win as your probability. You can weight closer matches more heavily.

There is no training step, which makes kNN easy to understand. It also makes it fragile: the answer depends heavily on which features you use, how you scale them, and how many neighbours you pick.

The maths

d(a,b)=∑f(af−bfsf)2d(a, b) = \sqrt{ \sum_{f} \left( \frac{a_f - b_f}{s_f} \right)^2 } p^=1k∑i=1kyi\hat{p} = \frac{1}{k} \sum_{i=1}^{k} y_i
  • d(a, b) is the distance between matches a and b.
  • a with subscript f is match a's value on feature f; s with subscript f is a scale for that feature so none dominates.
  • k is the number of neighbours; y with subscript i is 1 if neighbour i was a home win, 0 if not.
  • p-hat is the predicted probability.

In words: find the k most similar matches after putting features on a common scale, and average their results.

Worked betting example

Today's match: home side's xG difference over the last six games is +0.6, and the home implied probability is 50% (2.00 on Betfair). Scale xG by 0.5 and implied probability by 0.05.

Eight past matches (xG diff, implied probability, home win yes or no) and their distances:

  • 0.7, 49%, yes: distance 0.28
  • 0.5, 52%, yes: distance 0.45
  • 0.4, 47%, no: distance 0.72
  • 0.6, 45%, no: distance 1.00
  • 0.9, 55%, yes: distance 1.17
  • 0.1, 40%, no: distance 2.24
  • 1.2, 62%, yes: distance 2.68
  • 0.3, 35%, no: distance 3.06
  1. With k = 5 the nearest five are the first five rows: 3 home wins from 5, so p = 60%.
  2. A £10 back at 2.00 wins £10. EV = 0.60 × £10 − 0.40 × £10 = +£2.00, or +£1.88 after 2% commission.

Here is the problem. The standard error of a proportion from 5 matches is √(0.6 × 0.4 ÷ 5) ≈ 22 percentage points. The honest answer is "somewhere between about 17% and 100%", which tells you nothing. You need hundreds of neighbours for a usable estimate, and then they stop being very similar.

Where it's good

  • A quick sanity check: "how did matches like this actually go?"
  • Finding comparable matches for manual analysis.
  • Low-dimensional problems with lots of data, such as in-play states defined by score, time and pre-match price.
  • A nonparametric benchmark to test whether a fancier model adds anything.

Limitations and pitfalls

  • Small k gives wildly noisy probabilities; large k blurs everything towards the average. Choose k by out-of-sample log loss, not by eye.
  • Scaling decides the answer. Leave odds in raw units and xG in raw units and one will swamp the other.
  • The curse of dimensionality: with more than a handful of features, nothing is truly near anything. Consider dimensionality reduction first.
  • Leakage: make sure neighbours are only drawn from matches before the one you are predicting.
  • If market odds are a feature, kNN mostly echoes the market. Check it adds value beyond the price.
  • A logistic regression on the same features usually gives smoother, better-calibrated probabilities.

How to build it

  • Python: scikit-learn KNeighborsClassifier with StandardScaler; R: the class or kknn package.
  • Data: a few well-chosen pre-match features and results; more rows matter more than more columns.
  • Tip: use distance weighting and choose k with time-ordered cross-validation; report the neighbour count behind every prediction.
Learn it step by step
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members