Diego Garcia

live · review room

VAR

Every edge goes to review.

KairoX Labs

Three football models that check themselves against the bookmakers.

Tempo
1X2
Blaze
BTTS
Storm
O/U
Kick-off

What it is

Three independent models. One market each. Trained per league.

VAR puts probabilities on football matches in three markets: match outcome, both teams to score, and total goals. Each market has its own model, and each league its own champion per model.

The bookmakers' closing odds are not a feature. With the margin stripped out, they are the benchmark: every model is scored against the de-vigged closing line on the same matches.

There is no neural network anywhere. It is gradient-boosted trees plus Poisson rating models.

It exists to find out whether the models know anything the market doesn't. The checking is the project.

The name: variance, and the video assistant referee.

Three angles

Architecture

Pick a camera, pick a league. The diagram redraws from that league's champion.

championLa Liga · N 1,900 · ensemble · F 32 · w 0.53

Features

raw (N_raw, 101)

↓ 91 candidate features

what they are
  • 37 rolling
  • 13 running-table
  • 6 head-to-head
  • 4 attack/defence ELO
  • 25 style/matchup
  • 6 others

home-minus-away deltas, shift(1)-lagged per team

↓ select F = 32

X: (1,900, 32)

Branch A: 12 XGBoost members

4 regimes × 3 seeds

multi:softprob, num_class=3

4210422042

slow deep: lr 0.018 · 764 rounds · depth 6 · gamma base

(B, F) → (B, 3)

Uniform mean

stack (12, B, 3)

2,257 rounds × 3 seeds × 3 classes = 20,313 trees

(B, 3)

Branch B: Poisson member

Team ids

(B, 2)

Rates λ

λ_home = exp(μ + home_adv + att[h] − def[a])

λ_away = exp(μ + att[a] − def[h])

L-BFGS-B, time decay ξ = 0.002/day, no L2

2 + 2T + 1 scalars · T = 26

(B, 2)

Score grid

0 to 12 goals

Outlined: the 4 low-score cells with the Dixon-Coles correction. Then renormalise.

(B, 13, 13)

Read-out

triangle sums

(B, 3)

Log-opinion pool

p ∝ p_xgb^(1−w) · p_poisson^w

log-opinion pool, row-normalised. w is one scalar fitted on out-of-fold rows.

w = 0.53

XGBoostPoisson

(B, 3)

Calibration

softmax(log p / T)

temperature scaling, one scalar, bounds [0.25, 4.0]

(B, 3) = P(home), P(draw), P(away)

Under review

Training and promotion

A challenger has to win the replay, then clear every check.

Walk-forward CV · expanding by season · 3 folds

21-2222-2323-2424-2525-26
fold 0traintrainvalidate
fold 1traintraintrainvalidate
fold 2traintraintraintrainvalidate
finalrefit on all N rows. nothing held out.

Minimum two training seasons. A partial season is dropped before any of this starts.

Registry · step 1 · rank

  • Paired bootstrap on per-match out-of-fold log loss, against the champion.
  • 10,000 resamples.
  • The challenger must win in ≥ 95% of them.

10,000 resamples

≥ 95% to pass

Registry · step 2 · gates

check complete

Checking: Tempo

  • beats climatology
  • market gap not worse than the champion's
  • ECE ≤ 0.05
  • return CI upper bound > 0
  • draw recall > 0
  • fold std < 0.02, worst fold ≤ mean + 0.05

Checking: Blaze

  • beats climatology
  • market gap not worse than the champion's
  • AUC CI lower bound > 0.50
  • ECE ≤ 0.05
  • return CI upper bound > 0
  • BTTS rate in [0.40, 0.70]
  • fold std < 0.015, worst fold ≤ mean + 0.03

Checking: Storm

  • beats climatology at 2.5
  • no line worse than climatology + 0.005
  • monotonicity = 100%
  • ECE ≤ 0.05 on all lines
  • worst fold ≤ 0.03
  • ratings sanity

A league's first run is seeded as champion regardless of gates.

Offside lines

Shared vs separate

What the three models have in common, and what never crosses between them.

Shared

  • One raw CSV per league feeds all three models.
  • One partial-season rule and one walk-forward CV scheme. Duplicated code, same logic.
  • The same Poisson rate formula in all three.
  • The same four regimes and three seeds in Tempo and Blaze.
  • The same registry rule.

Never crosses

  • No fitted object crosses models. Each league fits three separate Poisson models.
  • Tempo's and Blaze's Poisson fits use identical inputs and settings, so their ratings are equal. Each still computes its own.
  • Feature caches are separate. Blaze recomputes its three Tempo-style quality features itself.
  • No cross-package import. All three carry their own copy of the training logic.
  • Tempo (B, 3)
  • Blaze (B,)
  • Storm (B, 5)

Slate / dashboard

The three outputs join in one place only: the slate and dashboard layer, per fixture.

Laws of the game

Principles

Why it is built this way.

  1. 1

    Walk-forward only

    Folds go by season, in order. Never shuffled. Every headline metric is pooled out-of-fold predictions.

  2. 2

    The closing line is the benchmark

    Closing odds are deliberately kept out of the features. De-vigged, they are what every model is scored against, on the same matches.

  3. 3

    Rank, then gate

    Candidates are ranked on log loss alone. Everything else is a gate, not a ranking criterion. This replaced a composite score that counted log loss and Brier as if they were two things.

  4. 4

    Pre-registered rules, tracked separately

    Rules written down in advance are kept apart from rules the grid search found.

The discipline, not any single model, is the product.

Team sheet

Stack

Six items.

Python
 
XGBoost
the tree members
SciPy / scikit-learn
 
Optuna
 
Streamlit + Plotly
dashboard
Scrapling
scraper
Replays

Related writing

Results, experiments and failures live in Briefs and Logs.