In one sentence
Quantile regression predicts a chosen percentile of an outcome, such as the minute by which the first goal has arrived in nine games out of ten, rather than only its average.
How it works
Linear regression predicts the average and then usually assumes a symmetric bell curve around it. The minute of the first goal, cricket totals and price moves are often lopsided: a goal in the first five minutes and a goalless 90 are not equally likely distances from the middle.
Quantile regression fits a separate line for each percentile you care about: the 10th, the median (50th), the 90th and so on. Each line can respond differently to the inputs. Two attacking sides might pull the 10th percentile much earlier while barely moving the 90th.
Stack several fitted percentiles together and you get a picture of the whole outcome range without assuming its shape. That is exactly what you need to price over/under lines away from the middle.
The maths
- τ: the percentile you want, as a fraction (0.9 for the 90th).
- Q_τ(y | x): the τ-th percentile of the outcome given inputs x.
- β^(τ): weights that can differ for each percentile.
- ρ_τ: the pinball loss, which the fit minimises; it punishes under-predictions by τ and over-predictions by 1 − τ.
In words: tilt the error penalty so that, for the 90th percentile, missing low costs nine times as much as missing high, and the best line ends up with 90% of outcomes below it.
Worked betting example
A first-goal timing model, fitted on past league matches, predicts the minute of the first goal tonight (illustrative figures). Goalless games are recorded as minute 95, and first-half stoppage time counts as minute 45. The model gives: 10th percentile minute 6, median minute 32, 90th percentile minute 82.
- The First Half Over/Under 0.5 market asks whether the first goal comes by minute 45. That sits between the median and the 90th percentile.
- Interpolating in a straight line: (45 − 32) ÷ (82 − 32) = 0.26 of the way from the 50th to the 90th. So about 50% + 0.26 × 40% ≈ 60.4% of outcomes fall at or before minute 45.
- P(first-half goal) ≈ 60.4%, P(no first-half goal) ≈ 39.6%. Fair prices: over 0.5 ≈ 1.66, under 0.5 ≈ 2.53.
- Betfair offers first-half over 0.5 at 1.80. £10 stake at 2% commission: a win pays £8 × 0.98 = £7.84. EV ≈ 0.604 × £7.84 − 0.396 × £10 ≈ +£0.78.
- Under 0.5 at 2.20 has EV ≈ −£1.38.
- Checking the pinball loss for the 90th percentile line (minute 82): if the first goal comes on 60, loss = 0.1 × 22 = 2.2; if the game ends goalless (95), loss = 0.9 × 13 = 11.7.
The straight-line interpolation is rough; fitting more percentiles (60th, 70th, 80th) near the line gives a better estimate.
Where it's good
- Football timing markets such as first-half goals, where the minute of the first goal is heavily skewed.
- Cricket runs markets, where totals are skewed and depend heavily on conditions.
- Player props: runs, points, passing yards, where the tails matter more than the average.
- Race margin or finishing-time bands in racing.
- Trading: estimating how far a price could move against you before exit (similar to value at risk).
- Any market where the line is set far from the average outcome.
Limitations and pitfalls
- Fitting each percentile separately can produce crossing lines (the 90th below the 80th); use methods that enforce order.
- Extreme percentiles (5th, 95th) need a lot of data; with a few hundred matches they are noisy.
- Interpolating between a few percentiles gives rough probabilities, especially in the tails.
- Weather and team news can shift the whole distribution late; stale inputs mean stale prices.
- Goalless games have no first-goal minute; recording them as minute 95 works only for percentiles below the goalless share.
How to build it
- Python: statsmodels QuantReg, scikit-learn QuantileRegressor, or lightgbm with objective set to quantile; R: quantreg.
- Data: past matches or outcomes with pre-match features (teams, expected goals, venue, weather).
- Fit a grid of percentiles, then check calibration: about 10% of outcomes should fall below the 10th percentile line.
- Tip: gradient-boosted quantile models handle interactions well but need careful validation to avoid overfitting.
Related methods
- Linear regression: models only the average.
- Normal distribution: the symmetric shape quantile regression avoids assuming.
- Log-normal distribution: a common shape for skewed totals.
- Value at risk: a percentile-based risk measure for trading.
- Extreme value theory: for the far tails where data is scarce.