The question
"If I build a model on my data and test it on the same data, I'm cheating. So how do I test it properly?"
Walk-forward testing is the answer used by professionals in betting and finance. It's how Statometrics judges every strategy: on data it never saw.
The idea in one sentence
Build the model using only the past, test it on the next slice of time it has never seen, then move the window forward and repeat, so every test bet is placed exactly as it would have been in real life.
The picture
Lay your seasons out in a row, oldest on the left. Walk-forward testing slides a window along them:
| Fold | Build on | Test on |
|---|---|---|
| 1 | 2018/19, 2019/20, 2020/21 | 2021/22 |
| 2 | 2019/20, 2020/21, 2021/22 | 2022/23 |
| 3 | 2020/21, 2021/22, 2022/23 | 2023/24 |
| 4 | 2021/22, 2022/23, 2023/24 | 2024/25 |
Each fold rebuilds the model from scratch: ratings, parameters, filters, the lot. It then bets through the test season without changing anything. Stitch the four test seasons together and you have four years of results from a model that never saw a single one of those matches in advance.
There are two ways to move the window:
- Rolling. A fixed-length build window that drops the oldest season as it adds the newest (the table above). Better when the game or the market changes over time.
- Anchored (expanding). The build window always starts at the first season and grows. Better when data is scarce and the patterns are stable.
Choose one before you start. Trying both and keeping the better one is another hidden test (Lesson 7.4).
Worked Betfair example
Your Match Odds model backs selections at average odds of 2.5, £10 a bet. Built and tested on all seven seasons at once, it shows +9.4% ROI. Now walk it forward. (Illustrative figures.)
- Break-even first. At 2.5 with 2% commission on winnings, break-even is 1 ÷ (1 + 1.5 × 0.98) = 40.5% winners.
- Fold 1, test 2021/22. 77 of 180 won (42.8%). Profit: 77 × £15 × 0.98 − 103 × £10 = +£101.90, an ROI of +5.7%.
- Fold 2, test 2022/23. 82 of 210 won (39.0%). Profit: −£74.60, an ROI of −3.6%.
- Fold 3, test 2023/24. 82 of 195 won (42.1%). Profit: +£75.40, an ROI of +3.9%.
- Fold 4, test 2024/25. 86 of 205 won (42.0%). Profit: +£74.20, an ROI of +3.6%.
- Pool by stake. 790 bets, £7,900 staked, +£176.90 profit: an out-of-sample ROI of +2.2%. That's about a quarter of the in-sample +9.4%.
- Is +2.2% evidence? A bet at 2.5 swings by about £1.23 per £1, so the standard error over 790 bets is about 4.4% and t ≈ 0.51. On profit alone, it proves nothing yet.
- Now check CLV. The average CLV in each test season was +2.4%, +1.9%, +2.2% and +2.0%, positive even in the losing season. Pooled, that's +2.1%. If one bet's CLV swings by about 12 points (Lesson 7.3), the standard error is 12 ÷ √790 ≈ 0.43 and t ≈ 4.9.
Verdict: the in-sample +9.4% was mostly fitted noise. Out of sample, the profit is small and not yet significant, but the model beat the closing price in every season, which is strong evidence of a real, smaller edge. That's a strategy worth running at modest stakes, with a CUSUM watching it.
Rules that keep it honest
- Everything is rebuilt inside the build window, including any tuning. If you choose a setting by looking at the test seasons, they're not test seasons any more (hyperparameter optimisation).
- Each test season is run once. Change the model after seeing it and that season joins the building data.
- Nothing flows backwards in time. Ratings carried into a test season must be built from matches before it.
- Prices and commission as in Lesson 7.5: the price on offer at bet time, 2% on winnings.
- Live betting is the last fold. After the final test, freeze the model and keep measuring. Live results are simply the next walk-forward season.
The formula
Pooled out-of-sample ROI
- K is the number of test folds.
- Profit_k and Staked_k are the profit after commission and the total staked in test fold k.
In plain English: add up all the test-season profit and divide by all the test-season stakes. Don't average the fold percentages, which over-weights small seasons.
The shrinkage ratio
- ROI_OOS is the pooled walk-forward ROI.
- ROI_in-sample is the ROI when the model is built and tested on the same data.
In plain English: how much of the backtest survived contact with unseen data. Here it's 2.2 ÷ 9.4 ≈ 24%. Close to 100% suggests a simple, robust model; close to zero or negative means the in-sample result was mostly noise.
Significance of the pooled result
- σ is the swing of one bet per £1 staked (or of one bet's CLV).
- n is the total number of test bets across all folds.
In plain English: the same test as Lesson 1.1, applied only to bets the model never saw. Run it on CLV as well as profit; CLV gets there far sooner.
Try it
Pen and paper: three test seasons show +£40 on £1,000 staked, −£25 on £1,500 and +£30 on £1,000. What's the pooled out-of-sample ROI, and how does it compare with averaging the three ROIs?
Answer: pooled is +£45 on £3,500 = +1.29%. Averaging +4.0%, −1.67% and +3.0% gives +1.78%, which overstates it because the losing season had the most bets.
Common mistakes
- Shuffling matches into a random split. Football is a time series. A random split lets the model learn from the future.
- Tuning on the test seasons. "I tried five settings and kept the one that walked forward best" turns the test into a building period.
- Quoting the best fold. One good season out of four is not the result. The pooled figure is.
- Expecting in-sample numbers live. Plan staking around the out-of-sample figure, and expect live results to be smaller still.
- Forgetting what was known when. Walk-forward only protects you if each fold also follows the no-look-ahead rule for prices, line-ups and ratings.
Testing on data the model never saw is the first line of the checklist in why good models still lose money.