In one sentence
Anomaly detection measures how unusual each new observation is compared with normal behaviour, and raises a flag when it crosses a threshold.
How it works
Every experienced trader has a sense of what a normal market looks like ten minutes before kick-off. Anomaly detection turns that feel into numbers: typical volume, typical price movement, typical spread.
The simplest version is a z-score: how many standard deviations the latest value sits from its usual average. More advanced methods, such as isolation forests, score how easy a point is to separate from the rest without assuming any shape for the data.
The hard part is not spotting anomalies, it is dealing with the false alarms. Scan enough markets and "rare" events happen every day by chance.
The maths
- x is the new observation, such as traded volume in a 5-minute window.
- μ (mu) is the normal average and σ (sigma) the normal standard deviation.
- N is how many observations you check; z-star is your alert threshold.
- P(Z above z-star) is the chance of a normal observation exceeding it.
In words: measure distance from normal in standard-deviation units, and remember that checking many windows multiplies the false alarms.
Worked betting example
In football Match Odds markets before kick-off, a team's traded volume per 5-minute window usually averages £2,000 with a standard deviation of £600 (illustrative). This window it trades £5,000 while the price shortens from 8.0 to 5.5.
- z = (5,000 − 2,000) ÷ 600 = 5.0. That looks extreme.
Now count the checks. You monitor 50 matches a day with 60 windows each: 3,000 checks.
- Under a normal distribution, the chance of z above 3 is about 0.135%, giving 3,000 × 0.00135 ≈ 4 false alarms a day at a threshold of 3.
- Real volumes have fat tails. Using a t-distribution with 3 degrees of freedom (same average and spread), the chance of exceeding 3 standard deviations is about 0.69%, giving roughly 21 false alarms a day.
So a z of 5 is notable, but a threshold of 3 would bury you in alerts. And spotting the steamer does not mean backing it: at 5.5 the price may already have moved past fair value.
Where it's good
- Flagging steamers and drifters for manual review in pre-match trading.
- Monitoring your own bot for runaway behaviour: unusual stake sizes, repeated failed orders, strange fills.
- Cleaning data: catching bad prices, wrong results or duplicated rows before they poison a model.
- Integrity work: unusual volume in low-profile markets can indicate informed money.
Limitations and pitfalls
- "Normal" changes: volume rises sharply towards kick-off, for big matches and at weekends. Use a baseline that matches the time and market type.
- Fat tails make normal-based thresholds far too sensitive; use robust measures like the median absolute deviation or extreme value theory.
- Scanning thousands of windows guarantees false alarms. Account for it, as in multiple testing corrections.
- An anomaly is not a trade signal. Following steamers after they are flagged often means buying late at a worse price.
- If you set thresholds using the same period you evaluate on, your hit rate will look better than it will be.
- Integrity flags need care. Unusual is not proof of wrongdoing, and naming names on thin evidence is unwise.
How to build it
- Python: rolling z-scores in pandas; scikit-learn IsolationForest or LocalOutlierFactor; pyod for more options.
- Data: Betfair stream or historical price data with traded volume by time bucket.
- Tip: build baselines per time-to-kick-off bucket and market type, and track how many alerts fire per day before trading on any of them.
Related methods
- Change-point detection finds when behaviour shifts for good, not just a single spike.
- Normal distribution underlies the z-score and its limits.
- Extreme value theory models the tails properly.
- Weight of money and VWAP give context to volume spikes.
- Multiple testing corrections control the false alarm rate.