In one sentence
Hyperparameter optimisation is the search for the best settings of a model or strategy, such as an Elo K-factor, a minimum edge threshold or a Kelly fraction, by testing many combinations and keeping the one that scores best.
How it works
Every model has dials that are set by you rather than learned from the data: how fast ratings update, how far back to look, which odds range to bet in. Hyperparameter optimisation turns those dials automatically and measures the result each time.
The common search methods are grid search (try every combination on a grid), random search (try random combinations, often more efficient) and Bayesian optimisation (use past results to decide where to try next). Tools such as Optuna make this quick.
The danger is that the more settings you try, the more likely the "best" one is just the luckiest one. In betting, where results are noisy and edges small, this is the main way people fool themselves.
The maths
There is no single formula; it is a search. The target is:
- θ is one combination of settings.
- Θ is the set of all combinations you allow.
- Score is the measure you optimise, such as log loss of the model or ROI of the strategy, on data not used to fit the model.
- θ* is the chosen best combination.
In plain English: try many settings, score each on held-back data, keep the best, then check it again on fresh data you have never touched.
Worked betting example
We built a deliberately useless model: 3,000 bets at random odds between 1.50 and 10.00, with the market prices fair and the model's "probabilities" equal to the market's plus random noise. There is no edge at all. Stakes are £10 and 2% commission is charged on winnings, so the true expected return is slightly negative, about −1.5%.
We tuned two settings on the first 2,000 bets: the minimum model edge to bet (0% to 10%) and the maximum odds (3.0 up to 10.0). That is 66 combinations, keeping only those with at least 100 bets.
- Best combination in-sample: bet when the model edge is over 3% and odds are 5.00 or less.
- In-sample: 290 bets, ROI +12.6%, profit £366.49.
- Out-of-sample on the last 1,000 bets: 152 bets, ROI −2.5%, a loss of £38.62.
We repeated the whole experiment 300 times with fresh random data:
- The best in-sample ROI averaged +13.0%, and was positive 93% of the time.
- The same settings out of sample averaged −1.2%, and were positive only 47% of the time, a coin flip.
A 13% backtest ROI from tuning 66 settings on noise is normal, not impressive.
Where it's good
- Tuning a model's structure (K-factors, decay rates, regularisation strength) using a proper score like log loss.
- Setting a small number of strategy rules when you have thousands of out-of-sample bets to check them on.
- Comparing staking fractions with simulations rather than past profits.
- Speeding up machine learning work, where models such as gradient boosting have many dials.
Limitations and pitfalls
- Overfitting is the default outcome. Every extra combination tried raises the best in-sample score by luck alone.
- Tuning directly on profit or ROI is especially risky, because profit is much noisier than log loss or Brier score.
- Reusing the same test set after peeking turns it into training data. Keep a final hold-out you only use once.
- Sports change: settings tuned on past seasons can drift as markets and rules change.
- Small samples in narrow filters (for example "odds 3.0 to 3.5, edge over 7%") produce big, meaningless ROIs.
- Market prices already reflect a lot of information, so a tuned strategy that beats the closing line on paper deserves extra suspicion.
How to build it
- Optuna, scikit-learn's GridSearchCV and RandomizedSearchCV, or hyperopt.
- Use walk-forward validation: tune on earlier seasons, test on the next, roll forward.
- Optimise a probability score (log loss) for the model, then judge betting rules separately on a final hold-out.
- Practical tip: count and record every combination you try, and judge the winner against what the same search finds on shuffled or random data.
Related methods
- Walk-forward validation – the right way to test tuned settings over time.
- Backtest bias checks – catching the self-deception this page warns about.
- Multiple testing corrections – adjusting for many tries.
- Genetic algorithms – another search method with the same overfitting risk.
- Regularisation – a common hyperparameter that itself fights overfitting.