Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Staking, bankroll and portfolio

Genetic Algorithms

A search method that breeds and mutates candidate strategies over many generations, powerful for messy problems but prone to overfitting.

Intermediateevaluationtradingstaking

In one sentence

A genetic algorithm searches for good strategy settings by imitating evolution: it keeps a population of candidate strategies, lets the fittest ones "breed", adds random mutations, and repeats for many generations.

How it works

Each candidate strategy is written as a short list of settings, called a genome. For a betting rule it might be the minimum edge, the odds range and the Kelly fraction. Each genome is scored on historical data; that score is its fitness.

The fitter genomes are more likely to be chosen as parents. Two parents swap parts of their settings (crossover) to make children, and a few settings are nudged at random (mutation) so the search keeps exploring. Over dozens of generations the population drifts towards settings that score well.

It suits messy scoring with thresholds and jumps that ordinary optimisers cannot handle. The flip side is that it is extremely good at fitting past noise.

The maths

There is no single formula; it is a search procedure. The common "roulette wheel" selection rule is:

P(select i)=Fi∑j=1NFjP(\text{select } i) = \frac{F_i}{\sum_{j=1}^{N} F_j}
  • F with subscript i is the fitness score of candidate i (it must be positive for this rule).
  • N is the population size.
  • The bottom sum adds up the fitness of every candidate.

In plain English: each candidate's chance of becoming a parent is its share of the population's total fitness. Tournament selection (pick a few at random and keep the best) is a common alternative that also works with negative scores.

Worked betting example

A population of four staking rules, each written as minimum edge, maximum odds and Kelly fraction, with fitness equal to bank growth in a validation period:

Rule Min edge Max odds Kelly fraction Fitness Chance of selection
A 2% 8.00 0.50 8 8 ÷ 40 = 20%
B 5% 4.00 0.50 12 12 ÷ 40 = 30%
C 1% 10.00 1.00 4 4 ÷ 40 = 10%
D 3% 6.00 0.25 16 16 ÷ 40 = 40%
  1. The wheel picks D and B as parents.
  2. Crossover after the first setting: child 1 takes D's minimum edge and B's other settings, giving 3%, 4.00, 0.50. Child 2 gets 5%, 6.00, 0.25.
  3. Mutation nudges child 1's Kelly fraction from 0.50 to 0.40.
  4. The children are scored and the next generation begins.

Now the warning. We ran a genetic algorithm with 50 rules for 40 generations, 2,000 evaluations, searching minimum edge and an odds range on 2,000 bets from a model with no edge whatsoever (fair prices, 2% commission, at least 100 bets required). Across 100 repeats, the best rule found averaged +45.9% ROI in-sample. On the next 1,000 bets the same rules averaged −1.1%, and were in profit only 45% of the time.

Where it's good

  • Searching rule-based trading strategies with many interacting thresholds, such as entry and exit triggers.
  • Problems where the score jumps around and cannot be smoothly optimised.
  • Tuning simulation-based staking plans, where each evaluation is a Monte Carlo run rather than past profit.

Limitations and pitfalls

  • Overfitting is the main risk, as the example shows: thousands of evaluations will always find something that looks brilliant on past data.
  • Optimising for profit or ROI chases noise; narrow rules with few bets and lucky results rise to the top.
  • Results depend on random seeds, population size and mutation rates, which are themselves settings to tune.
  • It can be slow: each fitness evaluation may be a full backtest.
  • Evolved rules often look sensible but have no reason to work; if you cannot explain why a rule should have an edge in an efficient market, treat it as luck.

How to build it

  • DEAP or PyGAD in Python; the ga package in R.
  • Set a minimum number of bets per rule, include 2% commission, and prefer smooth scores such as log growth over raw ROI.
  • Split data three ways: search on one part, choose between finalists on a second, and test once on a third.
  • Practical tip: rerun the whole search on shuffled or random data; if it finds similar "edges" there, the real result means nothing.
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members