Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Bayesian methods and time series

Kalman Filter

A recursive filter that separates a runner's underlying price from noisy trades, updating its estimate and its confidence every tick.

Advancedin-playtradingpre-match

In one sentence

The Kalman filter tracks a hidden quantity that drifts over time, such as a team's "true" win probability, by blending its last estimate with each new noisy observation according to how reliable each one is.

How it works

Last traded prices on Betfair bounce around. A single matched bet can print a price a few ticks away from where the market really sits. You want to know where the price actually is, not where the last £2 got matched.

The Kalman filter keeps two numbers: its current best estimate and how uncertain it is. Each step it first predicts: the estimate stays put but uncertainty grows, because the true price may have drifted. Then it updates: a new trade arrives and the filter moves part of the way towards it.

How far it moves is set by the Kalman gain. If the filter is unsure and trades are reliable, it moves a lot; if the filter is confident and trades are noisy, it barely moves. It is Bayesian updating for something that keeps changing.

The maths

For a simple "local level" model, where the true value wanders and each observation is noisy:

Pt∣t−1=Pt−1+Q,Kt=Pt∣t−1Pt∣t−1+RP_{t|t-1} = P_{t-1} + Q, \qquad K_t = \frac{P_{t|t-1}}{P_{t|t-1} + R} x^t=x^t−1+Kt(zt−x^t−1),Pt=(1−Kt)Pt∣t−1\hat{x}_t = \hat{x}_{t-1} + K_t (z_t - \hat{x}_{t-1}), \qquad P_t = (1 - K_t) P_{t|t-1}
  • x̂ₜ: the filtered estimate of the true implied probability.
  • zₜ: the new observation, here 1 ÷ last traded price.
  • P: the variance (uncertainty) of the estimate.
  • Q: how much the true value can drift per step.
  • R: how noisy each observation is.
  • Kₜ: the Kalman gain, the fraction of the surprise you accept.

In plain English: new estimate = old estimate plus a fraction of the gap to the latest trade, where the fraction depends on relative trust.

Worked betting example

The home side has been trading around 2.50 in Match Odds before kick-off, implied 40.0%. You work in implied probability because it behaves better than raw odds.

  1. Settings. Start with estimate 0.400, uncertainty P = 0.0004 (a standard deviation of 2 percentage points). Drift Q = 0.0001 per step; trade noise R = 0.0009.
  2. Trade at 2.30 (implied 0.4348). Predicted uncertainty = 0.0004 + 0.0001 = 0.0005. Gain K = 0.0005 ÷ 0.0014 ≈ 0.357. New estimate = 0.400 + 0.357 × 0.0348 ≈ 0.4124, price ≈ 2.42. Uncertainty falls to ≈ 0.00032.
  3. Trade at 2.34 (implied 0.4274). Predicted uncertainty ≈ 0.00042, gain ≈ 0.319. New estimate ≈ 0.4172, price ≈ 2.40.
  4. Trading read. Two trades printed at 2.30 and 2.34, but the filter puts the underlying price at about 2.40. It thinks those prints overstate the move.
  5. Does a lay at 2.34 pay? Laying £20 at 2.34 (liability £26.80), using 41.7% as the true chance: expected profit ≈ 0.583 × £20 × 0.98 minus 0.417 × £26.80 ≈ +£0.25. That is a thin edge on £26.80 of risk.

The filter is useful here mainly for stopping you chasing noise. The small edge it shows is only as good as the Q and R settings behind it.

Where it's good

  • Cleaning up last-traded-price series before feeding them into other signals.
  • Tracking team strength through a season as a slowly drifting rating.
  • Combining several noisy sources (Betfair price, bookmaker prices, your model) into one estimate, each with its own noise level.
  • In-play tracking of a scoring rate that changes with game state.
  • Setting how quickly a trading bot's "fair price" should react to new trades.

Limitations and pitfalls

  • It assumes smooth drift and normally distributed noise. Goals, red cards, team news and big gambles are jumps, which the basic filter smears out over several steps.
  • Q and R must be chosen or estimated. Set Q too small and the filter lags badly behind real moves; too big and it just copies every trade.
  • Last traded price is a noisy observation; best back and lay prices or weight of money may be better inputs.
  • Odds are bounded and skewed; filter in implied probability or log-odds, not raw decimal odds.
  • A filtered price is not a forecast of where the price goes next. It tells you where it is, not where it will be.
  • Small gaps between filter and market rarely survive commission.

How to build it

  • pykalman or filterpy in Python; statsmodels UnobservedComponents can estimate Q and R from data.
  • Data: time-stamped traded prices, or best back/lay prices, from the Betfair stream API or historical data.
  • Practical tip: estimate Q and R on past matches of the same type, then keep them fixed; re-tuning on every match overfits.
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members