Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Machine learning and simulation

Recurrent Networks and LSTMs

Neural networks with a memory that read sequences, such as price ticks or in-play match events, one step at a time.

Advancedin-playtrading

In one sentence

A recurrent network reads a sequence step by step, carrying a running memory forward, so its prediction depends on what happened before as well as now.

How it works

Watching a Betfair ladder, you do not judge the latest tick alone: you remember whether the price has been steadily shortening or bouncing around. A recurrent neural network (RNN) does the same by keeping a hidden state, a set of numbers summarising the sequence so far.

At each step the network combines the new input with its previous hidden state to produce an updated one. A basic RNN forgets older steps quickly. An LSTM (long short-term memory) adds gates, small learned switches that decide what to keep, what to forget and what to output, so it can hold on to information for longer.

In betting this is used for in-play event sequences and for exchange price series. It is also one of the most overhyped tools in the space: most short-term price movement is close to random, and an LSTM will find patterns in noise if you let it.

The maths

ht=tanh⁡(w xt+u ht−1+b)h_t = \tanh(w \, x_t + u \, h_{t-1} + b)
  • h with subscript t is the hidden state (the memory) after step t.
  • x with subscript t is the new input at step t, for example the price change in ticks.
  • w, u and b are learned weights: how much to weigh the new input, the old memory, and a constant.
  • tanh squashes the result to between −1 and 1.

In words: the new memory is a squashed blend of what just happened and what the network remembered. An LSTM wraps this in gates but the idea is the same.

Worked betting example

A toy RNN with w = 0.5, u = 0.8, b = 0 reads the last three one-minute price moves on the home team in pre-match Match Odds, in ticks: +2, −1, +3 (positive means shortening).

  1. Step 1: h = tanh(0.5 × 2 + 0) ≈ 0.762.
  2. Step 2: h = tanh(0.5 × −1 + 0.8 × 0.762) ≈ 0.109.
  3. Step 3: h = tanh(0.5 × 3 + 0.8 × 0.109) ≈ 0.920.
  4. An output layer turns this into a 70.7% chance the price shortens another tick.

You back £50 at 5.0 planning to lay one tick lower at 4.9 (ticks are 0.1 apart between 4 and 6).

  1. If it shortens: lay stake = £50 × 5.0 ÷ 4.9 ≈ £51.02, locking in £1.02, or £1.00 after 2% commission.
  2. If it drifts one tick to 5.1 and you cut: lay stake ≈ £49.02, locking in a loss of about £0.98.
  3. Break-even hit rate ≈ 49.5%. In backtesting the model's 70% signals shortened only 55% of the time, giving EV = 0.55 × £1.00 − 0.45 × £0.98 ≈ +£0.11 per trade.

That 55% versus the model's 70% is typical: raw sequence-model outputs are overconfident, and the real-world figure is what matters.

Where it's good

  • In-play modelling from event streams: possession and shot sequences in football, point-by-point tennis.
  • Summarising a pre-match price history into features for another model.
  • Very large tick datasets where there is enough data to learn temporal structure.
  • Multi-step forecasting, such as projecting a football side's in-play shot and xG rate over the next ten minutes.

Limitations and pitfalls

  • Short-term exchange prices are close to a random walk. A model that learns tick patterns from one period often finds nothing in the next.
  • Leakage through time is easy: normalising a price series with future values, or labelling with a move that overlaps the input window, produces fake accuracy.
  • Your backtest assumes you get filled at 5.0 and 4.9. Model the fill properly with queue position and your own reaction time, or the backtest counts trades you would not have got.
  • A simple momentum rule or order book imbalance feature often captures most of what an LSTM finds, with far less overfitting risk.
  • Transformers have largely replaced LSTMs for long sequences, though LSTMs remain lighter to run.

How to build it

  • Python: PyTorch (nn.LSTM, nn.GRU) or Keras. Keep networks small: one layer, 16-64 units to start.
  • Data: time-stamped Betfair stream data or event feeds, split by date into train, validation and a final untouched test period.
  • Tip: compare against a naive baseline (the price will not move) and a simple momentum rule before believing any LSTM result.
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members