In one sentence
A support vector machine (SVM) finds the boundary that separates two outcomes with the widest possible gap, then scores new matches by which side they fall on and how far.
How it works
Plot past football matches on a chart, say the home side's recent xG difference against its shots-on-target difference, and colour them by whether the home side won. An SVM draws the line that leaves the biggest margin between the two colours, paying most attention to the awkward matches near the border, called support vectors.
With a kernel trick, the SVM can bend that line into a curve without you building extra features. A penalty setting controls how many matches are allowed on the wrong side, trading tidiness for flexibility.
The catch for betting is that an SVM outputs a distance from the boundary, not a probability. You must convert it, usually with Platt scaling, before you can compare it with odds.
The maths
- w and b define the boundary; a smaller size of w means a wider margin.
- C is the penalty for matches on the wrong side of the margin.
- y with subscript i is +1 or −1 for the outcome; x with subscript i is the match's features.
- f(x) is the SVM's score for a new match; A and B are Platt scaling constants fitted on held-out data.
In words: find the widest-margin boundary while penalising mistakes, then map distance from that boundary to a probability.
Worked betting example
An SVM rates a home favourite with a score f(x) = 0.8 (illustrative). Platt scaling, fitted on a separate season, gives A = −1.5 and B = 0.1.
- Probability = 1 ÷ (1 + e to the power (−1.5 × 0.8 + 0.1)) = 1 ÷ (1 + e to the power −1.1) ≈ 75.0%.
- Betfair offers 1.40 on the home win in Match Odds, implying 1 ÷ 1.40 ≈ 71.4%.
- A £50 back wins £20. EV = 0.7503 × £20 − 0.2497 × £50 ≈ +£2.52 before commission.
- With 2% commission the win is £19.60, so EV ≈ 0.7503 × £19.60 − 0.2497 × £50 ≈ +£2.22.
Compare a weaker signal: a score of 0.2 converts to only about 55.0%. Without Platt scaling you would have no way to say whether a score of 0.8 is 60% or 90%, and staking on the raw score would be guesswork.
Where it's good
- Classification with a moderate number of features and a clear boundary, such as filtering which matches are worth modelling further.
- Small to medium datasets, where its margin-based penalty resists some overfitting.
- Situations where you only need a ranking (which of these bets looks best), not a precise probability.
Limitations and pitfalls
- Not a probability model. Platt scaling needs its own held-out data; fitting it on training data gives overconfident numbers. See Platt and isotonic scaling.
- Sensitive to feature scaling and to the choice of C and kernel width. Tuning these on the test set is a quiet form of leakage.
- Slow on large datasets: training time grows fast beyond tens of thousands of matches.
- Rarely beats logistic regression for betting probabilities, which gives calibrated outputs directly and is easier to explain.
- On big tabular datasets, gradient boosting is usually more accurate.
- Treat an SVM as a historical tool you may see in older papers. There are few betting problems where it is the best choice today.
How to build it
- Python: scikit-learn SVC with probability output or CalibratedClassifierCV; R: e1071 or kernlab.
- Data: scaled pre-match features, outcomes, and a separate time period for calibration.
- Tip: always compare calibrated SVM log loss against a plain logistic regression on the same later-season test set.
Related methods
- Logistic regression gives probabilities directly and is the benchmark.
- Platt and isotonic scaling turn SVM scores into probabilities.
- Calibration checks those probabilities hold up.
- Regularisation is the same idea as the C penalty.
- Gradient boosting usually wins on larger datasets.