In one sentence
A permutation test asks how often a random selection of your bets would perform as well as the ones your system picked, using reshuffling instead of textbook formulas.
How it works
Suppose your filter picks some bets out of a larger pool and those picks did better. The question is whether the filter found something or just got lucky. If the filter is useless, the label "picked" or "not picked" is effectively random.
So you test exactly that. Shuffle the labels at random, keeping the same number of picked bets, and recalculate the difference in profit. Do this thousands of times. The share of shuffles that match or beat your real difference is the p-value.
The appeal for betting is that it makes very few assumptions. Long odds, lopsided profit distributions and small samples all break the normal approximation, but a permutation test works with the real numbers you have.
The maths
- D obs: the observed difference, such as average profit of picked bets minus average profit of the rest.
- D star: the same difference after randomly reshuffling the picked and not-picked labels.
- The hash sign means "number of".
In plain English: the p-value is the fraction of random reshuffles that do at least as well as your real filter.
Worked betting example
A tiny example to show the mechanics. You had eight £10 football bets. Your filter picked four; the other four were in the pool but not picked. Profits before commission:
- Picked: won at 3.0 (+£20), won at 2.5 (+£15), lost (−£10), won at 2.2 (+£12). Total +£37, average +£9.25.
- Not picked: lost (−£10), lost (−£10), won at 1.8 (+£8), lost (−£10). Total −£22, average −£5.50.
Observed difference: 9.25 − (−5.50) = £14.75 per bet.
With only eight bets, you can list every way to choose four as "picked": there are 70. Checking all 70, only 5 give a difference of £14.75 or more (the real split is one of them).
p = 5 ÷ 70 ≈ 0.071.
So even a filter that looks dramatic on this sample would be matched by random labelling about 7% of the time. With real data you would have hundreds of bets, too many splits to list, so you draw 10,000 random shuffles instead and count the same way.
Where it's good
- Testing whether a filter (league, price band, day of week, kick-off time) adds anything over the full pool of bets.
- Longshot-heavy strategies, such as Correct Score or big-priced outsiders, where profit distributions are very skewed.
- Comparing two strategies on the same events by shuffling which strategy each result belongs to.
- Correcting for data mining: shuffle, rerun your whole search, record the best result, and see how often random data finds something as good as your best find.
Limitations and pitfalls
- Shuffling assumes bets are exchangeable. If bets cluster by date, league round or match, shuffle within those blocks, or the test will be too optimistic.
- It tests against the pool you give it. If the whole pool loses money, a filter that loses less will look "significant" and still lose.
- Commission and realistic prices must already be in the profits.
- With few bets, p-values come in coarse steps; with eight bets the smallest possible value is 1 ÷ 70.
- Thousands of shuffles on large datasets can be slow, especially if the whole model must be refitted each time.
- Like any test, running it on many filters and reporting the best one needs a multiple testing correction.
How to build it
- Python:
scipy.stats.permutation_test, or a loop withnumpy.random.permutation. - Data: per-bet profit after commission and the label you want to test, plus any grouping such as match ID.
- Tip: use at least 10,000 shuffles for p-values near 0.05, and fix the random seed so results are repeatable.
Related methods
- Hypothesis testing - the formula-based version of the same question.
- Bootstrapping - resampling to get confidence intervals rather than p-values.
- Multiple testing corrections - needed when you test many filters.
- Monte Carlo simulation - the general idea of learning from random repetition.
- P-values and confidence intervals - how to read the result.