In one sentence
Conditional probability is the chance of an event once you know some other event has happened, written as the probability of A given B.
How it works
Probabilities change as information arrives. The chance of a home win before kick-off is one number. The chance of a home win when the home side is 1-0 up at half-time is a different, higher number. That second number is a conditional probability.
The idea is to shrink the world you are looking at. Instead of all matches, you only look at matches where the condition is true, and ask how often the outcome happened within that smaller group.
This is how in-play markets think. Every goal, red card or passing minute changes the condition, and the price moves to the new conditional probability.
The maths
- P(A given B): the chance of A once we know B has happened.
- P(A and B): the chance that both A and B happen.
- P(B): the chance of the condition B happening at all.
In plain English: out of all the times B happens, what share of them does A also happen?
Worked betting example
Suppose your database holds 1,000 league matches (illustrative figures). In 380 of them the home side led at half-time. In 300 of those 380, the home side went on to win.
- P(home leads at HT) = 380 ÷ 1,000 = 0.38.
- P(home leads at HT and home wins) = 300 ÷ 1,000 = 0.30.
- P(home wins given home leads at HT) = 0.30 ÷ 0.38 = 0.789, so 78.9%.
- Fair odds for the home win at half-time in that situation = 1 ÷ 0.789 = 1.27.
If Betfair is offering 1.35 on a home side leading 1-0 at the break, and your conditional estimate is 78.9%, the EV per £1 is 0.7895 × 1.35 − 1 = 0.066, about 6.6% before commission. After 2% commission on the winnings it is 0.7895 × 0.35 × 0.98 − 0.2105 = 0.060, about 6.0%.
Before you bet, though, ask whether this match looks like the average match in your sample (see pitfalls).
Where it's good
- In-play pricing: goals, red cards and the minute all change the condition.
- Building simple in-play reference tables from historical data.
- Same-game combinations, where one outcome makes another more or less likely.
- Second-half goals: the chance of a goal after the break given it was 0-0 at half-time.
- Filtering systems: "home teams priced 1.80 to 2.00 in the first month of the season" is a conditional probability.
Limitations and pitfalls
- Small samples. The more conditions you stack, the fewer matches remain, and the estimate becomes noise. 12 matches is not a probability.
- Averages hide context. A strong favourite leading 1-0 is not the same as a weak underdog leading 1-0; condition on pre-match strength too.
- Direction matters. The chance of A given B is not the chance of B given A. Mixing them up is a common and costly error.
- Data mining risk. Search enough conditions and some will look profitable by pure luck.
- Old data may not reflect current game states, rules or tactics.
- The market already knows the obvious conditions. A half-time table of leads is priced in; value comes from conditions the market weighs poorly.
How to build it
- pandas groupby and crosstab are all you need for counting; statsmodels for smoothing with a regression once conditions multiply.
- Data: event-level match data with timestamps (goals, cards, minutes) plus pre-match prices to condition on team strength.
- Practical tip: always report the sample size next to every conditional probability, and do not act on groups below a few hundred.
Related methods
- Bayes' theorem: flips a conditional probability round the other way.
- Independence and correlation: when conditioning on B changes nothing, A and B are independent.
- Implied probability: the in-play price is the market's conditional probability.
- Markov chains: models where the next state depends only on the current one.