In one sentence
The multinomial distribution gives the probability of each combination of counts when every trial falls into one of several categories with fixed probabilities.
How it works
The binomial handles win or lose. Most betting markets have more outcomes: home, draw or away; 0, 1, 2 or 3+ first-half goals; one of several correct scores. The multinomial handles all of them at once.
Imagine rolling a loaded three-sided die 200 times, with sides marked H, D and A. The multinomial tells you how likely any particular split is, such as 88 home wins, 62 draws and 50 away wins. More usefully, it tells you whether a split is surprising given the probabilities you expected.
That makes it a natural tool for checking a model. If your model says 26% draws but the data shows 31%, is that luck or a flaw? A chi-square test built on the multinomial answers that.
The maths
- n is the total number of trials (matches); m is the number of categories.
- n1, n2 ... are the counts in each category; p1, p2 ... are their probabilities.
- O is the observed count and E the expected count (n × p) for each category.
- χ² (chi-square) measures how far observed counts are from expected ones.
In words: the first formula gives the chance of an exact split; the second summarises how unusual the split is overall.
Worked betting example
Match Odds, illustrative figures. Your model gives the same average probabilities across 200 matches: home 46%, draw 26%, away 28%. The results were 88 home wins, 62 draws and 50 away wins.
- Expected counts: 200 × 0.46 = 92, 200 × 0.26 = 52, 200 × 0.28 = 56.
- Chi-square = (88 − 92)² ÷ 92 + (62 − 52)² ÷ 52 + (50 − 56)² ÷ 56 = 0.174 + 1.923 + 0.643 = 2.74.
- With two degrees of freedom (three categories minus one), the p-value is 0.254.
- So ten extra draws over 200 matches is well within normal luck. It is not evidence your draw pricing is wrong.
- The exact probability of this precise split is only about 0.11%, but that is true of almost every specific split, which is why you test the overall distance instead.
Before backing more draws on the strength of this sample, you would need far more matches.
Where it's good
- Checking whether a 1X2 model's outcome frequencies match reality.
- Checking goal-band frequencies (0-1, 2-3, 4+ goals) against an Over/Under model.
- Simulating a batch of match outcomes for a season or accumulator.
- As the likelihood behind multinomial logistic regression for 1X2 probabilities.
Limitations and pitfalls
- Real matches have different probabilities each; the plain multinomial assumes identical ones. Group matches by predicted probability (calibration buckets) instead.
- Chi-square needs reasonable expected counts, roughly five or more per category; correct score cells often fail this.
- Passing a frequency test does not mean the model is profitable. The market may be just as well calibrated.
- It ignores the order of outcomes; for home, draw, away, a ranked probability score respects that draw sits between the other two.
- Small samples have little power: you will rarely detect a real two-point bias in draws with a few hundred matches.
How to build it
- scipy.stats.multinomial for exact probabilities; scipy.stats.chisquare for the test.
- numpy.random.multinomial to simulate outcomes.
- Data: model probabilities and results per match.
- Tip: test calibration by probability band, not just overall totals, which can hide offsetting errors.
Related methods
- Binomial: the two-outcome special case.
- Multinomial and ordinal logistic: regression models built on it.
- Hypothesis testing: the framework for the chi-square check.
- Ranked probability score: scoring ordered outcomes.
- Calibration: the more useful version of this check.