In one sentence
The AUC, short for area under the ROC curve, is the chance that your model gives a randomly chosen winner a higher score than a randomly chosen loser.
How it works
Pick a threshold, say "bet on anything the model rates above 40%". Some winners clear it (true positives) and some losers clear it too (false positives). Slide the threshold from 100% down to 0% and plot the true positive rate against the false positive rate: that plot is the ROC curve, short for receiver operating characteristic.
The area under that curve summarises it in one number. An AUC of 0.5 means the model ranks no better than a coin toss; 1.0 means every winner scored above every loser. Most sports models land somewhere between 0.6 and 0.8.
The key point for bettors is that AUC only cares about ranking. Multiply every probability by half and the AUC does not change at all, even though the probabilities are now useless for finding value.
The maths
- n W and n L: the number of winners and losers.
- s: the model's score or probability for each selection.
- The bracket counts 1 when the winner outscored the loser, and a half for a tie.
In plain English: compare every winner with every loser, and AUC is the share of those pairs your model ordered correctly.
Worked betting example
Your Match Odds model rated eight home sides to win their matches (illustrative figures). Four won and four did not.
- Winners' ratings: 0.70, 0.60, 0.45, 0.30.
- Losers' ratings: 0.50, 0.40, 0.25, 0.20.
There are 4 × 4 = 16 winner-loser pairs. Count how many the model got the right way round:
- Winner at 0.70 beats all four losers: 4.
- Winner at 0.60 beats all four: 4.
- Winner at 0.45 beats 0.40, 0.25 and 0.20: 3.
- Winner at 0.30 beats 0.25 and 0.20: 2.
Total 13 of 16, so AUC = 13 ÷ 16 = 0.8125.
That looks strong, but it says nothing about whether you should bet. If the winner rated 0.45 was trading at 2.00 (implied 50%), your model sees no value there despite ranking it well. And if the model's probabilities were all inflated by 10 percentage points, the AUC would still be 0.8125 while every value calculation would be wrong.
Where it's good
- Comparing feature sets during model building, such as whether adding shots data helps a football model separate winners from losers.
- Filters and classifiers where only ordering matters: flagging suspicious price moves, or choosing which matches to study.
- Spotting a model with no signal at all: an AUC near 0.5 out of sample is a quick red flag.
- Imbalanced problems, such as predicting rare events like red cards, where accuracy alone is misleading.
Limitations and pitfalls
- AUC ignores calibration completely, and value betting depends on calibration. Never judge a betting model on AUC alone.
- It gives equal weight to all parts of the ranking, but you usually bet only on the few selections where the model and the market disagree most.
- Pooling across events can mislead. In a three-way market, the right comparison is often between outcomes in the same match, not across all matches.
- A higher AUC than the market does not mean profit; the market's own AUC is often hard to beat, and commission still applies.
- Small samples give unstable AUCs; bootstrap a confidence interval.
- Tiny AUC gains from extra features often vanish out of sample.
How to build it
- Python:
sklearn.metrics.roc_auc_scoreandroc_curve; plot with matplotlib. - Data: model scores and binary results, preferably out of sample.
- Tip: compute the AUC of the margin-free market prices on the same events. If your model cannot match it, look elsewhere before tuning further.
Related methods
- Calibration - the missing half that AUC does not measure.
- Brier score - scores ranking and calibration together.
- Log loss - the usual training objective, sensitive to both.
- Logistic regression - a classic classifier often judged with AUC.
- Gradient boosting - high-AUC models that often need recalibration.