Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Machine learning and simulation

Clustering

Groups similar runners, teams or markets together without being told the answer, to reveal types and styles.

Intermediatepre-matchtradingevaluation

In one sentence

Clustering sorts items into groups of look-alikes, such as open, attacking teams versus deep-lying defensive ones, without you telling it what the groups should be.

How it works

Plot every team by how open its games are and how deep it defends. You will see clumps: some go toe-to-toe and concede at one end while scoring at the other, some sit in a low block and grind out tight games. Clustering finds those clumps automatically.

The most common method, k-means, picks k starting centres, assigns each item to the nearest centre, moves each centre to the average of its members, and repeats until nothing changes. Other methods find clusters of any shape or build a family tree of groups.

Clustering does not predict anything by itself. Its value is in creating labels, such as playing style or market type, that you then test as features in a predictive model.

The maths

min⁡∑j=1k∑x∈Sj∥x−μj∥2\min \sum_{j=1}^{k} \sum_{x \in S_j} \lVert x - \mu_j \rVert^2
  • k is the number of clusters you choose.
  • S with subscript j is the set of items in cluster j.
  • x is one item's features; μ with subscript j is the centre (average) of cluster j.
  • The double bars mean distance, squared.

In words: arrange items into k groups so that each item is as close as possible to the centre of its own group.

Worked betting example

Six teams, each with an openness score (1 means their games are the most end-to-end) and a deep-defending score (1 means they sit deepest). Illustrative figures:

  • A: 0.90, 0.30; B: 0.80, 0.40; C: 0.85, 0.20
  • D: 0.20, 0.90; E: 0.30, 0.80; F: 0.10, 0.70

Start k-means with centres at A and F.

  1. A, B, C are nearer to A's centre; D, E, F are nearer to F's.
  2. New centres: open teams (0.85, 0.30) and low-block teams (0.20, 0.80).
  3. Reassigning changes nothing, so the algorithm stops.

Suppose that in your own results, games between two teams from the open cluster went Over 2.5 goals 60% of the time over a decent sample, while the market priced Over 2.5 in those games at 1.90 (illustrative).

  1. Fair odds = 1 ÷ 0.60 ≈ 1.67.
  2. A £10 back on Over 2.5 at 1.90 wins £9: EV = 0.60 × £9 − 0.40 × £10 = +£1.40, or about +£1.29 after 2% commission.

The cluster is only the label; the 60% needs proper out-of-sample testing like any other angle.

Where it's good

  • Grouping teams by playing style (open, low-block, pressing, direct) to model style match-ups in Over/Under and Correct Score markets.
  • Creating running-style labels in racing when official data is missing or inconsistent.
  • Classifying exchange markets by how prices behave (steady, steaming, volatile) for trading rules.
  • Spotting segments where your model performs differently, for targeted review.

Limitations and pitfalls

  • You choose k, and different choices give different stories. There is often no single right number of clusters.
  • Results depend on scaling and on the random starting centres. Run several starts and check the groups are stable.
  • Clusters found on all your data and then tested on the same data leak information. Fit clusters on the training period only.
  • Slicing results by cluster, then by league, then by kick-off time is a recipe for false discoveries. Apply multiple testing corrections.
  • K-means assumes round, similar-sized groups. Real playing styles blur into each other.
  • Obvious clusters are usually already in the price. Everyone knows open teams produce goals, and the Over/Under market reflects it.

How to build it

  • Python: scikit-learn KMeans, DBSCAN or GaussianMixture; R: cluster or mclust.
  • Data: a small set of well-scaled descriptive features per team, player or market.
  • Tip: use the silhouette score and your own judgement to choose k, then test whether the cluster label improves out-of-sample log loss.
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members