In one sentence
Reinforcement learning (RL) trains an agent to choose actions, such as hold, green up or cut, by rewarding it for good outcomes over many simulated attempts.
How it works
Think of a new trader learning when to exit a position. At first they act almost randomly; over time they remember which actions in which situations tended to pay off. RL does the same with a table or a neural network of action values.
The agent sees a state (price, time to kick-off, open position), picks an action, gets a reward (profit or loss), and moves to a new state. After enough episodes it learns which action is worth most in each state, including actions that only pay off later.
The trap is where those episodes come from. You cannot afford millions of real trades, so the agent learns inside a simulator or a replay of historical data, and the simulator almost always makes fills and price reactions look kinder than reality.
The maths
- Q(s, a) is the estimated long-run value of taking action a in state s.
- α (alpha) is the learning rate: how far to move towards the new evidence.
- r is the immediate reward; γ (gamma) is how much future rewards count compared with now.
- s′ is the next state and the max picks the best action available there.
In words: nudge your estimate of an action's value towards what you just got plus the best you expect from here.
Worked betting example
A pre-match trading agent has backed the home team in Match Odds for £50 at 4.0. In the state "price one tick shorter at 3.95, 3 minutes to kick-off", its current value for "hold" is Q = 0.5 (in £, illustrative).
- It holds. Over the next step the price ticks back, and its marked-to-market reward is r = −£0.20.
- The best value available in the new state is £1.00 (say from a later green-up opportunity).
- With α = 0.1 and γ = 0.9: new Q = 0.5 + 0.1 × (−0.20 + 0.9 × 1.00 − 0.5) = 0.5 + 0.1 × 0.20 = 0.52.
So "hold" in that state is now valued slightly higher, £0.52. Over millions of updates the table settles into a policy. Every one of those numbers depends on the simulator's assumptions, so the policy must be checked against live results.
Where it's good
- Sequential decisions where today's action changes tomorrow's options: when to green up, when to add to a position, how to scale stakes through a session.
- Problems with a trustworthy simulator, such as optimising in-play decisions against a well-tested match model.
- Research and intuition-building: discovering exit rules you might not have thought of, then testing them simply.
- Games with fixed rules, where RL has had its famous successes.
Limitations and pitfalls
- The simulator gap: RL agents exploit any flaw in the environment. If your replay assumes instant fills at the displayed price, the agent will learn to scalp a profit that does not exist.
- Enormous overfitting risk: the agent memorises the historical price paths it trained on. Test on unseen days, and expect much worse results.
- Rewards are noisy and delayed, so learning is slow and unstable. Results can change a lot with the random seed.
- Commission, bet delays in-play and suspensions must be in the reward, or the policy is fantasy.
- For most retail traders, simple rules plus dynamic programming on a small model beat RL for effort versus reward.
How to build it
- Python: stable-baselines3 with a custom gymnasium environment; or a hand-written Q-table for small state spaces.
- Data: historical Betfair stream data with full ladder depth, plus a realistic fill model and commission.
- Tip: start with a tiny state space you can inspect by hand, and paper-trade live before trusting any learned policy.
Related methods
- Dynamic programming solves the same problem exactly when you know the model.
- Markov chains describe the state transitions RL learns about.
- Agent-based models can provide a richer simulated market.
- Market impact is what most RL simulators leave out.
- Backtest bias checks are essential before believing any RL result.