Altora Analytics

METHODOLOGY

How we build, validate and score our models: the data, the methods, the thresholds, and how to check our record yourself. Everything except the secret sauce.

← the record  ·  altora home

The discipline

How every model has to earn its place

The same process applies to every sport. A model doesn't go live because it looks clever. It goes live because it survives being tested on data it has never seen.

The models

What's actually under the bonnet

Different sports need different machinery. Here's what each one runs on: the methods and the feature families, if not the weights.

⚽ Football

Gradient boosting + Poisson

CatBoost (classification and regression) for match outcomes, plus a Poisson goal model driving Over/Under and Correct Score. Probabilities are Platt-calibrated.

Around 90 engineered features across 24 modules, including:

rolling xGopponent-adjusted strengthElo ratingsschedule difficultymarket-implied strengthhome-advantage indexrest daysshot tempovenue-specific formvolatilitymatch importanceleague environment

Markets: 1X2 · Over/Under 2.5 · BTTS · Correct Score · Accumulators · Misc.

🏀 Basketball

MOV-Elo + CatBoost

A margin-of-victory Elo with adaptive K and idle-decay (a team's rating settles as it plays, and drifts back toward average when idle), fed into CatBoost with league and country context.

rating differencerolling scoring marginopponent-adjusted formhome/away splitrest & back-to-backs
170,243games modelled
93,225out-of-sample tests
70.1%OOS accuracy

Totals & handicap (validated 27 Jul 2026): a separate regression CatBoost predicts each game's total points and winning margin from the same Elo and rolling-scoring features plus the league's scoring environment. Uncertainty (σ) is estimated from held-out residuals — per league where the sample allows — and converted to over/under and handicap-cover probabilities via the normal distribution. It beat a rolling-league-average baseline in both out-of-sample halves (margin error −19%, totals −6–7%). Currently recording bookmaker opening lines silently; no public record until the rules are pre-registered.

Trained pre-2024, tested 2024 onward. Calibrated at every band.

🤾 Handball

Two-stage 3-way

Handball draws (5.7% of games), so a 2-way model would be wrong. We run a two-stage model: a decisive-winner model multiplied by a separate draw-probability model, giving a true home/draw/away distribution.

MOV-Elo tuned specifically for handball, plus CatBoost with league, country and gender context. Men's and women's teams are kept strictly separate.

Totals & handicap (validated 27 Jul 2026): the same regression approach as basketball — CatBoost predicts each game's total goals and winning margin, with uncertainty (σ) estimated from held-out residuals and converted to over/under and handicap-cover probabilities via the normal distribution. It beat a rolling-league-average baseline in both out-of-sample halves (margin error −28–29%, totals −6.5–8%). Recording bookmaker opening lines silently; no public record until the rules are pre-registered.

126,968games modelled
76.5%winner accuracy
5.7%draw rate handled

🎯 Match picks

Darts · Snooker · Rugby · Esports

A shared adaptive-K Elo engine with idle-decay and, where it validates, margin-of-victory scaling, so a convincing win moves a rating more than a narrow one. Esports additionally runs dedicated per-title rating models for Dota 2 (OpenDota pro-match data) and CS2 (PandaScore map-level data) where they beat the shared pool on held-out data, with the detail in the honesty section below.

Since 22 July 2026, published probabilities blend the Elo rating with the market's opening price through a gradient-boosted calibration layer, validated out-of-sample per sport before switching, and only kept where it beat the Elo-only model on held-out data. Rugby Union remains on its own API-based ratings model.

Rugby covers Union and League with competition-aware ratings (70.7% Union / 65.1% League out-of-sample).

These are our free sports. Same discipline, simpler machinery, because the data supports one market: the winner.

📈 Systematic trading

FX · Indices · Commodities · rule-based, no discretion

Three independent systematic books (FX on 4h, Indices and Commodities on daily), each built from engineered market features and blended into one cross-asset portfolio. Fully rule-based, with no human overrides.

Validation is deliberately brutal: walk-forward testing, performance required to hold across separate market eras (not just one lucky regime), and stress-tested at double the real trading costs. Strategies that only work in one era, or die under cost stress, get cut.

The books are blended because they're near-uncorrelated with each other. That's where the portfolio's stability comes from, not from any single strategy.

From prediction to bet

How a pick is chosen, and our exact thresholds

A probability isn't a bet. A market can be entered by two avenues, value or probability, and each has a published bar it must clear. The bars are not secret, and they are not the edge: the model is. Every gate below was chosen on one half of held-out history, confirmed on the other, and frozen on 26 July 2026 for the whole 2026-27 season, with one pre-committed revision date of 1 November 2026. Markets where no gate cleared both halves are not backed at all, and we say which.

Market🔥 Value (EV) gate🎯 Probability gate
Match result (1X2)switched off on 29 July 2026: still published and settled, but never backed no gate confirmedprob ≥ 65% and odds ≥ 1.55
Over/Under 2.5EV ≥ 12%, prob ≥ 60%, odds ≥ 1.70prob ≥ 65% and odds ≥ 1.55
Both teams to scoreEV ≥ 8%, prob ≥ 55%, odds ≥ 1.70prob ≥ 55% and odds ≥ 1.70
Over/Under 3.5one combined gate: EV ≥ 2%, prob ≥ 60%, odds ≥ 1.40
Over/Under 1.5held at its pre-sweep gate (EV ≥ 3%, prob ≥ 78%, odds ≥ 1.20) not sweep-validated, see below
Double Chancepicked by probability alone: prob ≥ 80%, odds ≥ 1.10 value avenue switched off, see below
Which of these gates are sweep-validated, and which are not. Four cleared the both-halves test in the July 2026 sweep: 1X2 by probability, Over/Under 2.5 both ways, both teams to score both ways, and Over/Under 3.5. Three did not, and are marked above. The 1X2 value avenue was switched off on 29 July 2026: every threshold we tested made money on the first half of the data and lost it on the second, so rather than keep an unvalidated bar able to back bets, it now backs nothing. Over/Under 1.5 is held for the data reason below. Double Chance selects on probability because choosing those bets on modelled edge simulated at a 6% loss, so that avenue is switched off rather than dressed up. All three still publish every selection and its profit and loss, so you can see what we chose not to back, and all three are re-swept on clean data on 1 November.
Why every value gate also carries a probability floor. Without one, an "edge" on a 12%-probability longshot at odds of 11.00 passes the EV test, and that is almost always model noise rather than value. So each market's value avenue requires a minimum chance as well as a minimum edge (55% to 60% depending on the market). You will never see us call an 11.00 shot a value bet.
Why Over/Under 1.5 was not re-swept. A batch of historical quotes for that market turned out to be mis-lined, with higher-line prices stored in the 1.5 column. Roughly half the priced history was affected, so any threshold swept on it would be tuned to bad data. The bad prices have been voided and flagged, the capture path now rejects them, and the market keeps its old gate until there is enough clean history to sweep on 1 November. Tuning on numbers we do not trust would look more scientific and be worth less.
We publish the bets we skip. Selections clearing the bar are marked PASS, and those are the ones we back. Selections the model flagged but that failed a threshold are marked FAIL and are still published, with their own P&L. You can see the whole funnel, and judge for yourself whether our filter earns its place. Most services only show you what they backed.

How we score ourselves

The rules of our own record

Check us, don't trust us

How to verify our record yourself

A record you can't audit is just a screenshot. Ours is designed to be checked by someone trying to catch us out.

The part nobody else prints

What we don't claim

We don't beat the closing line. On basketball our model correlates 0.84 with the sharp closing price but the market is still sharper by roughly 0.016 Brier. Same story in rugby and football. The closing line sees things a results-based model can't: late team news, injuries, rotation. We mirror it well; we don't beat it. Anyone telling you their model beats the close is selling something.
Esports: two titles fixed, and we'll say which. Betfair's esports history carries no game label, so multi-game organisations (a club with CS2, Dota and LoL rosters) shared a single rating, a limitation we called unfixable with that data. It was: the fix was new data. Dota 2 (OpenDota pro-match data) and CS2 (PandaScore, 89,000 series with map margins) now run dedicated per-title rating models, each validated against the shared pool on held-out data before going live. Both roughly halve our gap to the market on their title. We tested the same upgrade for LoL and it failed our validation gate, so LoL and everything else stay on the shared pool with the safety gate: any pick that disagrees with the market by more than 20 points is withheld rather than published. That gate can only fire where a market price exists at the time we log, so a small number of published picks on thinly-traded markets sit outside it — 30 of our roughly 500 records to date, which we would rather tell you than quietly leave implied. Faults we find in our own data are written up on the corrections page; nothing is ever deleted.
A tracked record is not a promise. Past calibration doesn't guarantee future profit. Margins, market efficiency, model drift and sample size all bite. Never stake money you can't afford to lose.

Where we draw the line

What stays private, and why

We've told you the model families, the feature categories, the data sources, the thresholds, the staking and the scoring rules. What we keep back is the part that would let someone simply copy the work:

That's the difference between showing our workings and handing over the answer sheet. Everything you need to judge whether the models are any good is published: the record, the calibration, the thresholds and the funnel.