Markets coveredMatch OddsCorrect ScoreOver / UnderFirst HalfSecond Half
Statometrics
Model Library · Machine learning and simulation

Transformers

Neural networks that use attention to decide which past events matter most, powering modern language models and sequence forecasting.

Advancedpre-matchin-playtrading

In one sentence

A transformer looks at a whole sequence at once and learns which earlier items deserve the most attention when making a prediction.

How it works

A recurrent network reads events one at a time and hopes its memory holds on. A transformer instead lets every item look directly at every other item and score how relevant it is, a mechanism called attention.

For a team, that might mean paying more attention to last week's match against similar opposition than to a cup tie three months ago. The weights are learned from data rather than set by you, and many attention heads run in parallel, each picking up a different pattern.

Transformers are the engine behind large language models, which is where most betting use comes from today: reading news, injury reports and social posts. As a direct predictor on match data they are heavy machinery that rarely beats simpler models without very large datasets.

The maths

A full transformer has no single tidy formula, but its core step, attention, does.

Attention(Q,K,V)=softmax(QKTd)V\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{Q K^{T}}{\sqrt{d}}\right) V
  • Q (queries) describes what the current prediction is looking for.
  • K (keys) describes what each past item offers; Q times K gives a relevance score for each item.
  • d is the size of these vectors; dividing by √d keeps the scores in a sensible range.
  • softmax turns the scores into weights that add to 1; V (values) is the information that gets averaged with those weights.

In words: score every past item for relevance, turn scores into percentages, and take a weighted average.

Worked betting example

A model estimates a team's expected goals for Saturday from its last three matches, with xG of 2.1, 1.4 and 0.8. The attention layer scores their relevance as 1.5, 0.5 and 0.2 (the first was against a similar opponent).

  1. Softmax weights: e to the power 1.5, 0.5 and 0.2, each divided by their total, give 61.0%, 22.4% and 16.6%.
  2. Weighted xG = 0.610 × 2.1 + 0.224 × 1.4 + 0.166 × 0.8 ≈ 1.73.
  3. A plain average would give 1.43.

Using a Poisson model for the team scoring at least once:

  1. With 1.73: probability = 1 − e to the power −1.73 ≈ 82.2%, fair odds ≈ 1.22.
  2. With 1.43: probability ≈ 76.1%, fair odds ≈ 1.31.

If the team-to-score market is 1.28, the attention view says back and the simple average says leave it. Which is right can only be settled by testing both on hundreds of past matches.

Where it's good

  • Reading text: team news, press conferences, injury reports, via pre-trained language models (see NLP and sentiment).
  • Long event sequences with rich context, such as full-match event data or tick streams, when you have millions of rows.
  • Learning player or team embeddings across many competitions at once.
  • Summarising unstructured information into features for a simpler model.

Limitations and pitfalls

  • Needs huge amounts of data. Training one from scratch on a few seasons of league results is almost guaranteed to overfit.
  • Expensive to train and run, and in-play latency matters: a slow model is a stale model.
  • Attention weights look like explanations but often are not reliable ones. Do not read too much into them.
  • Large language models can state wrong facts confidently. Never let one feed unverified team news or odds into a live staking system.
  • Leakage through pre-trained models: a language model trained on text published after your test period may already know results.
  • For most pre-match problems, an expected goals feed plus a Poisson or boosting model gets you as far with a fraction of the effort.

How to build it

  • Python: PyTorch or Hugging Face transformers; use pre-trained models and fine-tune rather than training from scratch.
  • Data: large, time-stamped sequences or text corpora, with a clean cut-off date before your test period.
  • Tip: use a transformer to create features (text summaries, embeddings), then feed them into a simpler, well-calibrated model.
18+ only. Educational content, not financial or betting advice. Past results do not guarantee future returns. If gambling stops being fun, get free, confidential help at BeGambleAware.org.
Members