How we build, validate and score our models: the data, the methods, the thresholds, and how to check our record yourself. Everything except the secret sauce.
The same process applies to every sport. A model doesn't go live because it looks clever. It goes live because it survives being tested on data it has never seen.
Out-of-sample by time. We train on the past and test on a later period the model never saw. No shuffling, no peeking forward.
Features added one at a time. Each candidate feature is kept only if it improves out-of-sample log-loss. Ones that don't earn their keep get dropped, and we've thrown plenty away.
Calibration, not just accuracy. When the model says 70%, it has to happen about 70% of the time. We check this band by band and publish the chart.
Compared against the market. Every model is scored against bookmaker/exchange prices: the toughest benchmark there is.
Threshold gating. A prediction is not a bet. Only selections clearing our published thresholds get backed. Thresholds come from out-of-sample sweeps: choose on one half of held-out history, confirm on the other, adopt nothing that isn't positive on both halves with at least 30 bets in each. Frozen for the 2026-27 season on 26 July 2026, before round one, with a single pre-committed revision date of 1 November 2026. The frozen gates: 1X2 probability picks need 65%+ at odds 1.55+ (no 1X2 value threshold we tested was profitable on both halves, so on 29 July 2026 that avenue was switched off entirely: those selections are still published and settled, but they are never backed); Over/Under 2.5 value picks need a 12%+ modelled edge at 1.70+, probability picks 65%+ at 1.55+; BTTS value picks need an 8%+ edge at 1.70+ (55%+ chance), probability picks 55%+ at 1.70+; Over/Under 3.5 value picks a 2%+ edge at 1.40+; Double Chance has no profitable value gate on the data so far, so its value avenue is switched off and it is picked on probability alone (80%+), flagged as unvalidated until the November re-sweep; Over/Under 1.5 gates are held at their old values because a batch of mis-lined historical quotes contaminated that market's sweep data — it gets re-swept on clean captures on 1 November 2026 rather than tuned on numbers we do not trust.
Automated retraining. Models refresh on a schedule (weekly) so ratings never go stale.
The models
What's actually under the bonnet
Different sports need different machinery. Here's what each one runs on: the methods and the feature families, if not the weights.
⚽ Football
Gradient boosting + Poisson
CatBoost (classification and regression) for match outcomes, plus a Poisson goal model driving Over/Under and Correct Score. Probabilities are Platt-calibrated.
Around 90 engineered features across 24 modules, including:
A margin-of-victory Elo with adaptive K and idle-decay (a team's rating settles as it plays, and drifts back toward average when idle), fed into CatBoost with league and country context.
Totals & handicap (validated 27 Jul 2026): a separate regression CatBoost predicts each game's total points and winning margin from the same Elo and rolling-scoring features plus the league's scoring environment. Uncertainty (σ) is estimated from held-out residuals — per league where the sample allows — and converted to over/under and handicap-cover probabilities via the normal distribution. It beat a rolling-league-average baseline in both out-of-sample halves (margin error −19%, totals −6–7%). Currently recording bookmaker opening lines silently; no public record until the rules are pre-registered.
Trained pre-2024, tested 2024 onward. Calibrated at every band.
🤾 Handball
Two-stage 3-way
Handball draws (5.7% of games), so a 2-way model would be wrong. We run a two-stage model: a decisive-winner model multiplied by a separate draw-probability model, giving a true home/draw/away distribution.
MOV-Elo tuned specifically for handball, plus CatBoost with league, country and gender context. Men's and women's teams are kept strictly separate.
Totals & handicap (validated 27 Jul 2026): the same regression approach as basketball — CatBoost predicts each game's total goals and winning margin, with uncertainty (σ) estimated from held-out residuals and converted to over/under and handicap-cover probabilities via the normal distribution. It beat a rolling-league-average baseline in both out-of-sample halves (margin error −28–29%, totals −6.5–8%). Recording bookmaker opening lines silently; no public record until the rules are pre-registered.
126,968games modelled
76.5%winner accuracy
5.7%draw rate handled
🎯 Match picks
Darts · Snooker · Rugby · Esports
A shared adaptive-K Elo engine with idle-decay and, where it validates, margin-of-victory scaling, so a convincing win moves a rating more than a narrow one. Esports additionally runs dedicated per-title rating models for Dota 2 (OpenDota pro-match data) and CS2 (PandaScore map-level data) where they beat the shared pool on held-out data, with the detail in the honesty section below.
Since 22 July 2026, published probabilities blend the Elo rating with the market's opening price through a gradient-boosted calibration layer, validated out-of-sample per sport before switching, and only kept where it beat the Elo-only model on held-out data. Rugby Union remains on its own API-based ratings model.
Rugby covers Union and League with competition-aware ratings (70.7% Union / 65.1% League out-of-sample).
These are our free sports. Same discipline, simpler machinery, because the data supports one market: the winner.
📈 Systematic trading
FX · Indices · Commodities · rule-based, no discretion
Three independent systematic books (FX on 4h, Indices and Commodities on daily), each built from engineered market features and blended into one cross-asset portfolio. Fully rule-based, with no human overrides.
Validation is deliberately brutal: walk-forward testing, performance required to hold across separate market eras (not just one lucky regime), and stress-tested at double the real trading costs. Strategies that only work in one era, or die under cost stress, get cut.
The books are blended because they're near-uncorrelated with each other. That's where the portfolio's stability comes from, not from any single strategy.
From prediction to bet
How a pick is chosen, and our exact thresholds
A probability isn't a bet. A market can be entered by two avenues, value or probability, and each has a published bar it must clear. The bars are not secret, and they are not the edge: the model is. Every gate below was chosen on one half of held-out history, confirmed on the other, and frozen on 26 July 2026 for the whole 2026-27 season, with one pre-committed revision date of 1 November 2026. Markets where no gate cleared both halves are not backed at all, and we say which.
Market
🔥 Value (EV) gate
🎯 Probability gate
Match result (1X2)
switched off on 29 July 2026: still published and settled, but never backedno gate confirmed
prob ≥ 65%andodds ≥ 1.55
Over/Under 2.5
EV ≥ 12%, prob ≥ 60%, odds ≥ 1.70
prob ≥ 65%andodds ≥ 1.55
Both teams to score
EV ≥ 8%, prob ≥ 55%, odds ≥ 1.70
prob ≥ 55%andodds ≥ 1.70
Over/Under 3.5
one combined gate: EV ≥ 2%, prob ≥ 60%, odds ≥ 1.40
Over/Under 1.5
held at its pre-sweep gate (EV ≥ 3%, prob ≥ 78%, odds ≥ 1.20) not sweep-validated, see below
Double Chance
picked by probability alone: prob ≥ 80%, odds ≥ 1.10value avenue switched off, see below
Which of these gates are sweep-validated, and which are not. Four cleared the both-halves test in the July 2026 sweep: 1X2 by probability, Over/Under 2.5 both ways, both teams to score both ways, and Over/Under 3.5. Three did not, and are marked above. The 1X2 value avenue was switched off on 29 July 2026: every threshold we tested made money on the first half of the data and lost it on the second, so rather than keep an unvalidated bar able to back bets, it now backs nothing. Over/Under 1.5 is held for the data reason below. Double Chance selects on probability because choosing those bets on modelled edge simulated at a 6% loss, so that avenue is switched off rather than dressed up. All three still publish every selection and its profit and loss, so you can see what we chose not to back, and all three are re-swept on clean data on 1 November.
Why every value gate also carries a probability floor. Without one, an "edge" on a 12%-probability longshot at odds of 11.00 passes the EV test, and that is almost always model noise rather than value. So each market's value avenue requires a minimum chance as well as a minimum edge (55% to 60% depending on the market). You will never see us call an 11.00 shot a value bet.
Why Over/Under 1.5 was not re-swept. A batch of historical quotes for that market turned out to be mis-lined, with higher-line prices stored in the 1.5 column. Roughly half the priced history was affected, so any threshold swept on it would be tuned to bad data. The bad prices have been voided and flagged, the capture path now rejects them, and the market keeps its old gate until there is enough clean history to sweep on 1 November. Tuning on numbers we do not trust would look more scientific and be worth less.
We publish the bets we skip. Selections clearing the bar are marked PASS, and those are the ones we back. Selections the model flagged but that failed a threshold are marked FAIL and are still published, with their own P&L. You can see the whole funnel, and judge for yourself whether our filter earns its place. Most services only show you what they backed.
How we score ourselves
The rules of our own record
Fixed, published stakes. Singles are £10 level; accumulators are tiered by design (Short Odds 3-fold £10, Mid Odds 3-fold £5, Big Odds 4-fold £3, Match Doubles £5) and every stake is recorded per bet in the ledger. From 5 August 2026 the gated markets settle every game on the daily board (earlier rows are the published top fives only, disclosed on the record page); the Telegram channels receive the top five per approach. The acca construction rules themselves (leg counts, odds bands, which markets feed each acca) were chosen by a walk-forward backtest over four seasons and are fixed in advance, not tuned to look good afterwards. Those rules still stand, including the small edge floor on the Big Odds four-fold: we removed that floor briefly on 29 July 2026 and restored it the same day, because the acca shapes were chosen by a walk-forward test that included it. No accumulator was published while it was off. No hindsight-flattering staking plans, no doubling up.
Priced at the opener. P&L is struck at the odds we flagged when the pick was published, not a better price we found later.
Every bet counts. Wins, losses, and markets that are in the red. Nothing is quietly dropped.
Calibration published. Predicted probability vs what actually happened, band by band.
Brier score vs the market. We show our score next to the market's, even when the market wins.
Closing-line value (CLV). We record the opening price on every pick, and the closing price wherever the exchange gives us one before the off. Coverage is honest rather than complete: as of July 2026 it runs from 99% of snooker picks down to 45% of darts and none at all in tennis, because our capture windows do not yet line up with every start time. We are widening them. If our picks consistently beat the closing line that is real evidence, and if they don't that shows too, but we will not quote a closing-line figure on a sport where we have not actually captured the closes.
Voids & postponements. Settled only on a real result. A voided or abandoned match stays on the record page and in the downloadable CSV, marked VOID, and is excluded from the accuracy figures because there is no result to score. Nothing is removed.
Data: API-Sports (football, rugby union, basketball & handball fixtures, results and odds) · Betfair Exchange (live market prices and results across all sports) · football-data.co.uk (historical football odds) · tennis-data.co.uk (historical tennis results and odds) · OpenDota (Dota 2 match data) · PandaScore free tier (CS2, LoL, Dota 2 & Valorant match histories and schedules). Each sport's record page credits its own sources.
Check us, don't trust us
How to verify our record yourself
A record you can't audit is just a screenshot. Ours is designed to be checked by someone trying to catch us out.
Picks are timestamped by a third party, not by us. Every pick is committed to a public GitHub repository before the event starts. Git's commit history is timestamped by GitHub and can't be quietly rewritten. You can look back through the history and see exactly what we said, and when we said it.
Full record as CSV. Every record page has a download of the complete history (date, selection, odds, stake, result, P&L), not just the last few.
Losing markets stay visible. If a market is in the red, it stays on the page in red.
The part nobody else prints
What we don't claim
We don't beat the closing line. On basketball our model correlates 0.84 with the sharp closing price but the market is still sharper by roughly 0.016 Brier. Same story in rugby and football. The closing line sees things a results-based model can't: late team news, injuries, rotation. We mirror it well; we don't beat it. Anyone telling you their model beats the close is selling something.
Esports: two titles fixed, and we'll say which. Betfair's esports history carries no game label, so multi-game organisations (a club with CS2, Dota and LoL rosters) shared a single rating, a limitation we called unfixable with that data. It was: the fix was new data. Dota 2 (OpenDota pro-match data) and CS2 (PandaScore, 89,000 series with map margins) now run dedicated per-title rating models, each validated against the shared pool on held-out data before going live. Both roughly halve our gap to the market on their title. We tested the same upgrade for LoL and it failed our validation gate, so LoL and everything else stay on the shared pool with the safety gate: any pick that disagrees with the market by more than 20 points is withheld rather than published. That gate can only fire where a market price exists at the time we log, so a small number of published picks on thinly-traded markets sit outside it — 30 of our roughly 500 records to date, which we would rather tell you than quietly leave implied. Faults we find in our own data are written up on the corrections page; nothing is ever deleted.
A tracked record is not a promise. Past calibration doesn't guarantee future profit. Margins, market efficiency, model drift and sample size all bite. Never stake money you can't afford to lose.
Where we draw the line
What stays private, and why
We've told you the model families, the feature categories, the data sources, the thresholds, the staking and the scoring rules. What we keep back is the part that would let someone simply copy the work:
Trained model weights and artefacts
Hyperparameters and tuning
The exact feature engineering recipe and transformations
The training and data pipelines
That's the difference between showing our workings and handing over the answer sheet. Everything you need to judge whether the models are any good is published: the record, the calibration, the thresholds and the funnel.