In one sentence
Logistic regression turns a weighted sum of inputs into a probability between 0 and 1 for a two-way outcome like win or lose.
How it works
Linear regression can predict any number, including impossible probabilities like 1.3. Logistic regression fixes this by working in log-odds, where any value is allowed, then squeezing the result through an S-shaped curve into the 0 to 1 range.
Each weight tells you how much the log-odds move for a one-unit change in an input. Positive weights push the probability up, negative ones pull it down. The model is fitted by choosing the weights that make the observed outcomes most probable.
It is the workhorse of betting models: simple, fast, hard to break, and it often outputs well-calibrated probabilities. Many far fancier models fail to beat it once tested fairly on unseen data.
The maths
- p: the probability of the event (for example, over 2.5 goals).
- p ÷ (1 − p): the odds of the event; the log of this is the log-odds.
- x_1 to x_k: the inputs, such as rating gap or rest days.
- β_0 to β_k: the weights learned from past data.
In words: add up the weighted inputs to get log-odds, then convert log-odds back into a probability.
Worked betting example
An over 2.5 goals model fitted on past league matches (illustrative weights): log-odds of over 2.5 = −1.20 + 0.60 × (combined expected goals for the two sides) − 0.30 × (local derby: 1 if yes).
- Today: combined expected goals is 3.1 and it is a local derby.
- Log-odds = −1.20 + 0.60 × 3.1 − 0.30 × 1 = 0.36.
- Probability = 1 ÷ (1 + e^(−0.36)) ≈ 58.9%. Fair price ≈ 1.70.
- Betfair offers 1.71 on over 2.5. £10 stake at 2% commission: a win pays £7.10 × 0.98 = £6.96. EV ≈ 0.589 × £6.96 − 0.411 × £10 ≈ −£0.01. No bet.
- The break-even price after commission is about 1.72. At 1.80, EV rises to about +£0.51.
This is typical: a price just above your fair odds is not value once commission is included.
Where it's good
- Yes/no football markets: over/under goals lines, a goal before half-time, both teams to score.
- Match-winner markets in two-outcome sports such as tennis, snooker and darts.
- Combining ratings (Elo, xG, speed figures) with context like rest, travel and weather.
- Recalibrating another model's output, known as Platt scaling.
- A benchmark: if a complex model cannot beat logistic regression out of sample, drop it.
Limitations and pitfalls
- Assumes each input shifts the log-odds by a fixed amount; curved or interacting effects need to be added by hand.
- Three-way markets (home, draw, away) need a multinomial or ordinal version, not two separate logistic models.
- With many correlated inputs the weights become unstable; regularisation helps.
- Perfect separation in small samples, where one input predicts every outcome, sends weights to infinity.
- Good log loss is not the same as profit; compare against closing prices, not just results.
- Rare events (a 50-1 outsider winning) give few data points, so tail probabilities are shaky.
How to build it
- Python: scikit-learn LogisticRegression (regularised by default), statsmodels Logit for p-values; R: glm with family = binomial.
- Data: one row per event with the outcome and pre-event features only.
- Evaluate with log loss and calibration plots using walk-forward validation.
- Tip: include the market's implied probability as a feature to see whether your other inputs add anything beyond it.
Related methods
- Linear regression: the version for numeric targets.
- Multinomial and ordinal logistic regression: for three-way markets.
- Regularisation: stabilises weights when features pile up.
- Calibration: checks the probabilities mean what they say.
- Bradley-Terry model: logistic regression on rating differences.