In one sentence
A neural network passes your inputs through layers of weighted sums and on-off switches, learning the weights that turn stats into a probability.
How it works
Picture a logistic regression, then stack several of them so the outputs of one layer feed the next. Each small unit, or neuron, adds up its inputs with weights, then applies a switch such as ReLU, which passes positive values and turns negatives into zero.
Training shows the network thousands of past matches and adjusts every weight a little to reduce its error, using a method called backpropagation. With enough layers and data, it can represent almost any relationship.
That flexibility is the problem as well as the promise. On typical betting datasets of a few thousand to a few hundred thousand rows, a neural network rarely beats gradient boosting or a tidy logistic regression, and it is far easier to overfit.
The maths
- x is the list of inputs (for example rating difference and form).
- W and b are the hidden layer's weights and offsets; h is the hidden layer's output after the ReLU switch.
- v and c are the output layer's weights and offset.
- p is the predicted probability.
In words: mix the inputs a few different ways, switch off any negative mixes, then feed what is left into a logistic curve.
Worked betting example
A tiny network predicts a home win from two inputs: rating difference 0.4 and away-form score 0.2.
- Hidden neuron 1 = 1.5 × 0.4 − 0.5 × 0.2 + 0.1 = 0.6. ReLU keeps it at 0.6.
- Hidden neuron 2 = −1.0 × 0.4 + 2.0 × 0.2 + 0.0 = 0.0. ReLU leaves 0.
- Output score = 1.2 × 0.6 + 0.8 × 0 − 0.3 = 0.42.
- Probability = 1 ÷ (1 + e to the power −0.42) ≈ 60.3%.
The home side is 1.80 in Match Odds on Betfair, implying about 55.6%.
- A £10 back wins £8. EV before commission = 0.603 × £8 − 0.397 × £10 ≈ +£0.86.
- With 2% commission the win is £7.84, so EV ≈ 0.603 × £7.84 − £3.97 ≈ +£0.76.
Real networks have thousands of weights, not a handful. The more weights, the more ways to fit noise, so that 60.3% needs a calibration check before it is trusted.
Where it's good
- Very large datasets: tick-level exchange data, minute-by-minute in-play football, event and tracking data.
- Unstructured inputs such as text, images or sequences, via specialised versions like recurrent networks.
- Learning embeddings (compact numeric fingerprints) for players or teams from many interactions.
- Multi-output problems, such as predicting a full scoreline grid at once.
Limitations and pitfalls
- Data hungry. A few seasons of one league is far too little; the network will memorise it.
- On tabular betting data, simple models often win. Always benchmark against logistic regression and gradient boosting on the same later-season test set.
- Results vary with random seeds and settings. If a different seed changes your ROI from +3% to −2%, you do not have an edge.
- Overconfident probabilities are common. Check calibration and use dropout, weight decay and early stopping.
- Leakage hides easily in feature pipelines, for example normalising inputs using statistics from the whole dataset including the test period.
- Hard to explain, which makes it hard to spot when the model has learned something silly.
How to build it
- Python: PyTorch, TensorFlow/Keras, or scikit-learn's MLPClassifier for small experiments.
- Data: scaled numeric features, time-ordered splits, and far more rows than you think you need.
- Tip: start with one small hidden layer, log loss as the objective, and early stopping on a held-out later season; only grow it if validation improves.
Related methods
- Logistic regression is a neural network with no hidden layer, and the benchmark to beat.
- Gradient boosting is usually stronger on tabular data.
- Recurrent networks and LSTMs handle price and event sequences.
- Regularisation is essential to stop overfitting.
- Calibration checks whether the outputs are usable as probabilities.