Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Module 7 · Lesson 7.5

Building an honest backtest

“What must my backtest include?”

The question

"My backtest says +10%. What must an honest betting backtest include before I believe it?"

A backtest replays history and asks what would have happened if you'd bet this way. It's the cheapest research tool you have and the easiest to fool yourself with, because almost every mistake pushes the result the same way: up.

The idea in one sentence

An honest backtest places each bet at the moment the rule would really have fired, at the price really on offer then, pays 2% commission on winnings, and uses nothing that wasn't known at that moment.

The picture

Think of an honest backtest as a replay where the future is hidden. At each bet, freeze the clock and ask three questions:

Question The honest answer The flattering mistake
What price could I take? The best back price on offer at that timestamp The best price seen all day, or the last traded price when it suits
What did I pay? 2% commission on net winnings in each market Nothing
What did I know? Only data timestamped before the bet Season averages, final line-ups, closing prices, later results

Every row on the right makes your backtest look better than real life. They also stack: a system can survive one of them and still die from all three together.

Worked Betfair example

Your Over 2.5 goals system backs matches where both teams average more than 2.8 goals a game. Over three seasons it finds 500 bets at £10. (Illustrative figures.)

  1. The first version. 262 of 500 bets won. It priced every bet at the best price seen that day, averaging 2.10, with no commission. Profit: 262 × £11.00 − 238 × £10 = +£502, an ROI of +10.0%.
  2. Fix the price. Re-run using the best back price on offer at the moment the rule fired, one hour before kick-off. That averages 2.02. Profit: 262 × £10.20 − £2,380 = +£292.40, or +5.8%. The best-of-day price was information from the future: you couldn't know when it would appear.
  3. Add 2% commission. Commission is 2% of the net winnings in each market, and each bet here is a different match. Winnings of 262 × £10.20 = £2,672.40 lose £53.45. Profit: +£238.95, or +4.8%.
  4. Remove the look-ahead. The "averages more than 2.8 goals" filter used each team's whole-season average, including matches played after the bet. Recompute it from prior games only and the system now picks 430 bets, of which 219 won (50.9%).
  5. The honest result. Profit: 219 × £10.20 × 0.98 − 211 × £10 = +£79.12 on £4,300 staked, an ROI of +1.8%.
  6. Is +1.8% real? Each bet at 2.02 swings by about £1.01 per £1, so over 430 bets the standard error is about 4.9%. That's t ≈ 0.38, which is nothing like evidence (Lesson 1.1).

Verdict: the same idea went from +10.0% to +1.8%, and from "brilliant" to "no evidence either way". The system hasn't changed. The test has just stopped lying.

The full checklist

Before you trust any backtest, tick every line:

  • Timestamps on everything. Every input and every price has a time, and the bet time comes after all of them.
  • The price on offer at bet time. Not the best of the day, not the closing price, not a price you'd have had to wait for.
  • Only prior games. Ratings, averages and form built only from matches finished before the bet. A season average in October uses August to October only.
  • Team news when it was known. If your model uses line-ups, your bet time must be after they were announced.
  • 2% commission on net winnings, per market.
  • Every bet counted. Void and postponed matches return the stake; they aren't quietly dropped, and neither are awkward losers.
  • Rules frozen before the test. The test period is run once. Change the rules and it becomes part of the building data (Lesson 7.4).
  • CLV recorded. Log the closing price for every bet so you can see whether you beat the market, not just whether you won (Lesson 2.4).

The formula

ROI after commission

ROI=∑winnersS (Ot−1)(1−c)  −  ∑losersS∑all betsS\text{ROI} = \frac{\sum_{\text{winners}} S\,(O_t - 1)(1 - c) \;-\; \sum_{\text{losers}} S}{\sum_{\text{all bets}} S}
  • S is the stake on each bet.
  • O_t is the price on offer at the bet's timestamp, not the best or the closing price.
  • c is Betfair's commission on net winnings, 0.02.

In plain English: take each winner's profit at the price you could really have had, knock 2% off it, subtract the losing stakes, and divide by everything you staked.

The no-look-ahead rule

xi  =  f(data with timestamp<ti)x_i \;=\; f\big(\text{data with timestamp} < t_i\big)
  • x_i is any input your rule uses for bet i: a rating, an average, a price.
  • t_i is the moment bet i would have been placed.
  • f is however you calculate that input.

In plain English: every number feeding a bet must be calculable from data that existed before the bet. If it couldn't have been known then, it can't be used.

Break-even with commission

pbreak-even=11+(Ot−1)(1−c)p_{\text{break-even}} = \frac{1}{1 + (O_t - 1)(1 - c)}
  • O_t is the price taken, and c is 0.02.

In plain English: at 2.02 you need to win just over 50.0% of the time to break even after commission. At the best-of-day 2.10 you'd have needed only 47.6%, which is exactly why the flattering price flatters.

Try it

Pen and paper: your backtest backed 100 selections at an average of 1.95 (the price on offer at bet time) and 55 won. Work out the ROI after 2% commission.

Answer: winnings 55 × 0.95 = 52.25, less 2% = 51.205, minus 45 losing stakes = +6.205 on 100 staked, so +6.2% (before commission it was +7.3%).

Common mistakes

  • Using closing prices as bet prices. If your rule fires an hour before kick-off, the closing price is from the future. It belongs in the CLV column, not the P&L.
  • Building ratings from the whole data set. A rating that has "seen" the season's results knows who finished top. Build it forward, one match at a time.
  • Dropping awkward bets. Removing "unusual" matches, abandoned games or data errors after seeing their results always removes more losers than winners.
  • Testing on the data you designed the rules on. Honest prices don't fix a rule that was fitted to that very period. Test on data it never saw (Lesson 7.6).
  • Stopping at a good backtest. A backtest is a filter, not a verdict. The final test is live bets, and the first thing to watch is CLV (Lesson 7.3).

Backtests that ignore the real world are pitfall 6 in why good models still lose money.

Check yourself

1. Your backtest backs every selection at the best price it reached during the day. What's wrong?
2. Your filter uses each team's goals-per-game average for the whole season. Why is that a problem for a bet in October?
3. Your system backed 100 selections at 1.95 and 55 won. What's the ROI after 2% commission on winnings?
Key takeaway

Every backtested bet must use the price available at that moment, pay 2% commission, and know nothing the world didn't know yet. Anything else is a story, not a test.

Go deeper in the Model Library
Next lesson
7.6 Walk-forward testing →
How do I test without cheating?
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
MembersHow to Build an Honest Betting Backtest: Real Prices, Commission and No Look-Ahead — Statometrics Academy