Three model families feed one shared validation discipline. Nothing gets sized until the strategy layer agrees the edge is real.
Match result. Form and strength features feeding a calibrated classifier.
Both teams to score. Its own feature set, including a head to head variant built for this market alone.
Goal rate modelling. A time decayed Poisson layer runs alongside the tree ensemble, not inside it.
Every family draws from the same feature set per league: rolling form, table state, head to head history, Elo ratings, venue tendencies, and a quality gap between the two sides. Splits run strictly by season, never shuffled, and every out of fold prediction gets calibrated before anything downstream is allowed to trust it. Team and referee identities get reconciled across leagues before any of this touches a feature builder. Small detail, big difference two seasons later.
Five checkpoints. A candidate has to clear all of them before it ever touches a stake.
Compares calibrated output, raw output, and de-vigged market odds side by side. Edge has to beat the market, not just the raw model.
Shortlists thresholds and staking rules from three independent origins, then ranks them on seasons the search never saw.
Estimates the odds that a shortlisted strategy is a lucky in-sample fluke before it is allowed to advance.
Simulates thousands of bet sequences for risk of ruin, bankroll paths, and whether the edge beats a random baseline.
Sizes the position. The most aggressive plan whose risk of ruin still clears the bar, never flat Kelly.
The models are not the interesting part anymore. The interesting part is that nothing gets staked until three separate checks agree it should. That discipline, more than any single model, is the actual product.
Still testing. More leagues on the way.