Diego Garcia

Back to briefs

> kairox labs

One Number for Twelve Slips

Tempo prices 1X2, Blaze prices BTTS, Storm prices Over/Under 2.5, all on the same match. Multiply their calibrated probabilities into a same-match slip and something breaks completely: the predicted return is mathematically identical for all twelve resulting slips, whichever one you pick. Not a rounding error. A structural blind spot, and it comes with a price tag.

Aug 202613 min read

A cross-model slip is the obvious next step once you have three calibrated models on the same match: take Tempo's 1X2 leg, Blaze's BTTS leg, Storm's Over/Under 2.5 leg, multiply the three probabilities, and price a same-match parlay off the result. Twelve possible combinations per match. Before building that engine, this research asked the prerequisite question: can you actually multiply three model outputs together and trust the number that comes out?

The answer is no, measured across 5,329 matches in La Liga, Bundesliga and the Premier League (2021-22 through 2025-26, odds present for all three markets in 99.98% of matches). Not approximately no. The three markets are different functions of one scoreline, so multiplying their marginals asserts an independence the generating process forbids. The error this produces is large, has a predictable sign, and in its sharpest form removes the engine's ability to rank slips at all.

Three Slips That Cannot Exist

1X2, BTTS and Over/Under 2.5 are not three independent questions about a match. They are three different summaries of the same pair of integers: home goals and away goals. That makes some combinations logically impossible, and the algebra predicts exactly which ones before touching any data.

SlipWhy it cannot occurObserved / 5,329
H+YES+UNDERBTTS=YES needs both teams on the board; Under 2.5 with both scoring leaves only 1-1, a draw0
A+YES+UNDERSame argument, mirrored for an away win0
D+NO+OVERBTTS=NO and a draw forces 0-0, whose total is 0, which is Under0

Three of the twelve slips are dead by scoreline algebra. Zero observations confirms it in every league, individually and pooled.

A naive engine has no way to know this. Fed the de-vigged market as its three calibrated legs, it still assigns these three slips a combined 24.97% of its total probability budget (23.67% to 26.12% by league). A quarter of the engine's confidence sits on outcomes that cannot physically happen.

Naive multiplication misprices every slip, and prices three that cannot occur

Pooled, 5,329 matches. The three impossible slips have true probability exactly zero.

0.0%5.0%10.0%15.0%20.0%25.0%30.0%impossible,yet pricedimpossible,yet pricedimpossible,yet pricedH+YES+OVERH+YES+UNDERH+NO+OVERH+NO+UNDERD+YES+OVERD+YES+UNDERD+NO+OVERD+NO+UNDERA+YES+OVERA+YES+UNDERA+NO+OVERA+NO+UNDER
Empirical (true joint)Independence product (naive)
Figure 1: Empirical joint probability vs the independence product across all 12 slips, pooled. The three impossible slips (annotated) have zero true probability and non-zero naive probability.

A Constant Is Not a Ranking

The impossible slips are the loud failure. The quiet one is worse. Give the naive engine perfectly calibrated legs, the de-vigged market itself, and its predicted EV collapses to a single number for every slip on every match. Write the de-vigged leg probability as p = (1/o) / OR for odds o and market overround OR. Then for any slip:

naive_prob × parlay_odds = Π (pi × oi) = Π (1 / ORm) = 1 / (OR1x2 · ORbtts · ORou)

Every leg's odds and probability cancel to the inverse of its own overround. Multiply three legs and the slip structure disappears entirely: the result depends only on the three markets' vig, never on which of the twelve combinations you chose. With this dataset's mean overrounds (1X2 1.0299, BTTS 1.0529, O/U 1.0345, compounding to a 12.18% three-leg vig), the identity predicts −10.8545% for every slip. The realised mean of the same per-match quantity across 5,329 matches: −10.8437%. A naive engine built on calibrated marginals is not a weak ranking model. It cannot rank at all, which is a stronger and different claim than "inaccurate."

One identical EV for all 12 slips, but reality spans -100% to +79%

Realised parlay ROI per slip at naive-product prices, flat 1-unit stakes, 95% bootstrap CI (10,000 resamples, seed 42), pooled across 5,329 matches.

-125%-100%-75%-50%-25%0%25%50%75%100%125%+49%impossible-50%+60%-13%+79%impossible-2%+43%impossible-64%+69%H+YES+OVERH+YES+UNDERH+NO+OVERH+NO+UNDERD+YES+OVERD+YES+UNDERD+NO+OVERD+NO+UNDERA+YES+OVERA+YES+UNDERA+NO+OVERA+NO+UNDER-10.84%
underpriced by multiplicationoverpriced by multiplicationlogically impossiblenaive engine's predicted EV: -10.84% for all 12 slips
Figure 2: Realised ROI per slip against the naive engine's single flat prediction. This is not a live market price; same-game parlays are priced by bookmakers with their own correlation models. It measures the size of the dependence the naive engine ignores, not a bettable edge.

A Direction, Not Just a Size

The error also has a sign, and the sign is not arbitrary. Per-slip lift, true probability over the independence product, ranges from 0.53× to 1.93×. Every underpriced slip (lift above 1) describes a modal scoreline: 1-1 draws, tight wins, open games where both sides score. Every overpriced slip (lift below 1) describes an extreme one: a 3+ goal blowout with a clean sheet, or a high-scoring draw. Independence treats "who won" and "how many goals" as unrelated, when a clean sheet suppresses the total and a high total suppresses clean sheets. The sign of this gap matched the sign of lift − 1 in all twelve cells, including the one slip, D+NO+UNDER, whose lift (1.16×) sits close enough to the vig that its realised ROI is the only one not distinguishable from zero.

Multiplication is off by up to 1.9x, in both directions

Lift = true joint / independence product. Above the dashed parity line, multiplication underprices the slip; below it, multiplication overprices.

-0.25x0.00x0.25x0.50x0.75x1.00x1.25x1.50x1.75x2.00x2.25x1.59ximpossible — exactly 00.68x1.75x0.92x1.93ximpossible — exactly 01.16x1.62ximpossible — exactly 00.53x1.89xH+YES+OVERH+YES+UNDERH+NO+OVERH+NO+UNDERD+YES+OVERD+YES+UNDERD+NO+OVERD+NO+UNDERA+YES+OVERA+YES+UNDERA+NO+OVERA+NO+UNDER+1.00x
underpriced by multiplicationoverpriced by multiplicationlogically impossibleparity: independence would be correct
Figure 3: Lift per slip against parity. Above the line, multiplication underprices; below it, multiplication overprices. Impossible slips sit at exactly zero.

The pattern is consistent with which markets share the most information. Mutual information between BTTS and Over/Under is 0.2102 bits, the strongest pairwise term, because both are direct functions of the same two goal counts. 1X2 and BTTS share the least, 0.0453 bits: knowing who won says relatively little about whether both teams scored. Summed across all three markets, the total correlation, the KL divergence between the true joint and the independence product, is 0.5234 bits pooled, and every league lands in the same range (0.4857 to 0.5688). The effect is not a single-league artefact.

LeagueTotal correlation (bits)D+YES+UNDER ROID+YES+UNDER lift
La Liga0.5688+70.05% [+49.83, +91.22]1.895×
Bundesliga0.4857+93.73% [+63.95, +124.20]2.021×
Premier League0.5038+75.34% [+51.43, +99.99]1.926×

The headline 1-1 slip reproduces in every league. Bundesliga has the lowest total correlation of the three and the largest single-slip lift: dependence concentrates differently per league, it does not disappear.

The ROI figures are a measurement, not a strategy

The naive-product parlay price used here, the raw product of the three legs' own odds, is not a price any real bookmaker offers. Same-game parlays are precisely the product books price with their own correlation models, historically because naive-product pricing was exploitable this way. These ROI numbers express the size of the dependence in units that are legible (percentage points against a 12.18% vig), and establish that it is large enough to swamp that vig several times over. They are an upper bound on available edge, not a confirmed one, and this research does not claim a bettable edge exists.

What This Rules Out

Tempo, Blaze and Storm are not three independent models that happen to look at the same match. They are three separately-estimated marginals of one joint scoreline distribution, each trained with no knowledge that the other two exist. 1X2 is sign(h − a), BTTS is 1[h ≥ 1] ∧ 1[a ≥ 1], Over/Under 2.5 is 1[h + a ≥ 3]. Three functions of one random variable are dependent unless the functions happen to be orthogonal, and these read overlapping features of the same two integers by construction. Multiplying their outputs does not approximate the joint. It asserts a structure the generating process does not have, which is why the resulting bias is not small and not zero-mean: it has a known sign per cell and a fixed floor of impossible-outcome mass that no amount of model improvement removes, because it is not a model error.

The fix is architectural, not incremental, and it does not require new modelling. Storm's DecayPoisson already yields per-match λ_h, λ_a, which is a full scoreline distribution: 1X2, BTTS and Over/Under 2.5 all derive from it by summation. A slip engine built on that scoreline matrix is coherent by construction, impossible slips get exactly zero probability, not a patched exception, and every dependence measured here is reproduced for free. Tempo and Blaze's calibrated marginals still have a role, not as factors to multiply in, but as constraints the scoreline matrix should be fit to match, via minimum-divergence reconciliation rather than a product.

What is still open

This is a model-free result: no trained artefact for Tempo, Blaze or Storm exists in this tree, so it measures the outcome space and the market, not model skill. Two questions stay open until artefacts exist. Do the three models make correlated errors on the same match, over and above the outcome-space dependence measured here? And are they mutually coherent today, given that Tempo's P(draw) and Blaze's P(BTTS=NO) jointly imply a P(0-0) that may already contradict what Storm assigns to Under 2.5. A cheap per-match coherence audit, flagging exactly that kind of contradiction, is the more immediate deliverable and does not wait on any of it.