Altora Analytics
When the model says 70%, does it happen about 70% of the time? This page is the answer, measured on every settled prediction we have published.
Each point is a band of predictions. The vertical bars are 95% confidence intervals, so a short bar means plenty of data and a long bar means the band is still thin. Points sitting on the dashed line are perfectly calibrated. Sitting above it means the model is too cautious; below it means too confident.
| Model said | Predictions | Average stated | Actually won | Difference |
|---|---|---|---|---|
| 50 to 60% | 1223 | 54.7% | 56.9% | +2.2 pts |
| 60 to 70% | 1019 | 64.4% | 63.7% | -0.7 pts |
| 70 to 80% | 735 | 74.6% | 72.8% | -1.8 pts |
| 80 to 90% | 369 | 84.1% | 79.4% | -4.7 pts |
| 90 to 100% | 157 | 93.8% | 89.2% | -4.6 pts |
| Sport | Settled | Average stated | Actually won | Brier |
|---|---|---|---|---|
| Darts | 231 | 66.5% | 68.0% | 0.2051 |
| Snooker | 371 | 66.9% | 66.6% | 0.2102 |
| Rugby | 182 | 69.7% | 65.9% | 0.2173 |
| Esports | 260 | 65.3% | 65.8% | 0.2248 |
| Tennis | 2459 | 66.4% | 65.8% | 0.2165 |
Brier score measures both accuracy and honesty at once, and lower is better. Predicting 50% on everything scores 0.25, so anything below that is doing real work. It punishes confident wrong answers much harder than cautious ones.
Football's ledger tracks every selection its models flag. The overwhelming majority were rejected by our published thresholds and never backed; the 161 that cleared a threshold are judged in the PASS blocks on the record page, not here. This page also counts only selections with a stated probability, so its total sits below the record page's settled count. They sit apart from the chart above because a threshold-rejected pick's strike rate is not a measure of model performance. For completeness: 19723 settled selections, average stated probability 41.6%, actual strike rate 40.8%, Brier score 0.1795. The sample is far too small to conclude anything, and we would say so even if it flattered us.
Calibration is the claim we care most about, because it is the one that makes a probability useful rather than decorative. It is also the one most easily faked by people who quietly drop their losers, which is why every settled prediction we have ever published is in the numbers above, and why the raw data behind them is downloadable as CSV from every record page.
What it does not prove is profit. A perfectly calibrated model can still lose money if the prices are not there, and a badly calibrated one can get lucky for months. Those are separate questions, tracked separately on the record pages. The honest summary today is that the sample is young: bands with fewer than about fifty predictions should be read as indicative rather than settled, and the confidence bars in the chart show you exactly which ones those are.
Generated automatically from the live ledgers on 17 September 2026, so this page cannot drift from the record it describes. Voided and postponed events are excluded; nothing else is. Method: how the models are built and tested.