In one sentence
Dimensionality reduction replaces lots of overlapping statistics with a handful of combined scores that keep most of the information.
How it works
Shots, expected goals and touches in the box all measure roughly the same thing: how much a team attacks. Throwing all three into a model gives it three chances to fit noise and makes its coefficients unstable.
Principal component analysis (PCA) finds the combination of features that captures the most variation, then the next best combination at right angles to the first, and so on. Often the first one or two components hold most of the information.
You then model with those few components instead of the raw features. Other methods such as t-SNE and UMAP are popular for drawing pretty charts, but they are for visualising, not for feeding a betting model.
The maths
- x is a team's standardised features (each rescaled to average 0, spread 1).
- v with subscript 1 is the set of weights for the first component; z with subscript 1 is the resulting score.
- Var means variance, the spread of the scores across teams.
- λ with subscript j is the variance captured by component j.
In words: find the weighted blend of features that spreads teams out the most, and measure how much of the total variation it keeps.
Worked betting example
Six teams' per-match averages for shots, xG and box touches:
- Team 1: 16, 1.7, 30 · Team 2: 12, 1.5, 22 · Team 3: 10, 0.9, 21
- Team 4: 14, 1.4, 29 · Team 5: 9, 1.1, 15 · Team 6: 11, 1.0, 24
- Correlations: shots with xG 0.85, shots with touches 0.93, xG with touches 0.62. They overlap heavily.
- PCA on the standardised data gives a first component with weights 0.62, 0.55 and 0.57, a near-equal blend of all three.
- That first component explains 86.9% of the variation; the second 12.8%; the third 0.3%.
- Team scores on the first component: Team 1 +2.36, Team 4 +1.27, Team 2 +0.25, Team 6 −0.65, Team 3 −1.37, Team 5 −1.87.
One "attacking threat" number now does the job of three. In a goals model, you would estimate one coefficient instead of three wobbly ones. Six teams is far too few to trust in practice; the same steps apply to a full league over several seasons.
Where it's good
- Many correlated stats: team event data, player tracking metrics, sectional times.
- Reducing overfitting in small datasets by cutting the number of inputs.
- Preparing features for distance-based methods like k-nearest neighbours and clustering.
- Visualising a league or field to spot outliers.
Limitations and pitfalls
- PCA keeps the directions with most variation, not the ones that best predict results. The useful signal might sit in a small component you threw away.
- Fitting PCA on all seasons, then testing on one of them, leaks the test period's structure into training. Fit it on training data only.
- Components can flip sign or reshuffle between refits, which breaks live pipelines if you are not careful.
- Components are harder to explain than "shots per game".
- Regularisation such as ridge regression handles correlated features directly and often does as well without the extra step.
- t-SNE and UMAP distort distances; do not use their outputs as model inputs or read too much into their clusters.
How to build it
- Python: scikit-learn PCA in a Pipeline with StandardScaler; R: prcomp.
- Data: a table of numeric features per team, match or runner, standardised first.
- Tip: keep the fewest components that explain about 80-90% of variation, then check out-of-sample that the reduced model is no worse than the full one.
Related methods
- Linear regression often uses components as inputs.
- Regularisation is the alternative way to tame correlated features.
- Clustering works better on reduced features.
- K-nearest neighbours needs few dimensions to work well.
- Independence and correlation explains the overlap PCA exploits.