live · review room
Every edge goes to review.
Three football models that check themselves against the bookmakers.
Three independent models. One market each. Trained per league.
VAR puts probabilities on football matches in three markets: match outcome, both teams to score, and total goals. Each market has its own model, and each league its own champion per model.
The bookmakers' closing odds are not a feature. With the margin stripped out, they are the benchmark: every model is scored against the de-vigged closing line on the same matches.
There is no neural network anywhere. It is gradient-boosted trees plus Poisson rating models.
It exists to find out whether the models know anything the market doesn't. The checking is the project.
The name: variance, and the video assistant referee.
Pick a camera, pick a league. The diagram redraws from that league's champion.
championLa Liga · N 1,900 · ensemble · F 32 · w 0.53
Features
raw (N_raw, 101)
↓ 91 candidate features
home-minus-away deltas, shift(1)-lagged per team
↓ select F = 32
X: (1,900, 32)
Branch A: 12 XGBoost members
4 regimes × 3 seeds
multi:softprob, num_class=3
slow deep: lr 0.018 · 764 rounds · depth 6 · gamma base
(B, F) → (B, 3)
Uniform mean
stack (12, B, 3)
2,257 rounds × 3 seeds × 3 classes = 20,313 trees
(B, 3)
Branch B: Poisson member
Team ids
(B, 2)
Rates λ
λ_home = exp(μ + home_adv + att[h] − def[a])
λ_away = exp(μ + att[a] − def[h])
L-BFGS-B, time decay ξ = 0.002/day, no L2
2 + 2T + 1 scalars · T = 26
(B, 2)
Score grid
0 to 12 goals
Outlined: the 4 low-score cells with the Dixon-Coles correction. Then renormalise.
(B, 13, 13)
Read-out
triangle sums
(B, 3)
Log-opinion pool
p ∝ p_xgb^(1−w) · p_poisson^w
log-opinion pool, row-normalised. w is one scalar fitted on out-of-fold rows.
w = 0.53
(B, 3)
Calibration
softmax(log p / T)
temperature scaling, one scalar, bounds [0.25, 4.0]
(B, 3) = P(home), P(draw), P(away)
A challenger has to win the replay, then clear every check.
Walk-forward CV · expanding by season · 3 folds
Minimum two training seasons. A partial season is dropped before any of this starts.
Registry · step 1 · rank
10,000 resamples
≥ 95% to pass
Registry · step 2 · gates
check completeChecking: Tempo
Checking: Blaze
Checking: Storm
A league's first run is seeded as champion regardless of gates.
What the three models have in common, and what never crosses between them.
Shared
Never crosses
Slate / dashboard
The three outputs join in one place only: the slate and dashboard layer, per fixture.
Why it is built this way.
Folds go by season, in order. Never shuffled. Every headline metric is pooled out-of-fold predictions.
Closing odds are deliberately kept out of the features. De-vigged, they are what every model is scored against, on the same matches.
Candidates are ranked on log loss alone. Everything else is a gate, not a ranking criterion. This replaced a composite score that counted log loss and Brier as if they were two things.
Rules written down in advance are kept apart from rules the grid search found.
The discipline, not any single model, is the product.
Six items.
Results, experiments and failures live in Briefs and Logs.