Altora Analytics
When the model says 70%, does it happen about 70% of the time? This page is the answer, measured on every settled prediction we have published.
Each point is a band of predictions. The vertical bars are 95% confidence intervals, so a short bar means plenty of data and a long bar means the band is still thin. Points sitting on the dashed line are perfectly calibrated. Sitting above it means the model is too cautious; below it means too confident.
| Model said | Predictions | Average stated | Actually won | Difference |
|---|---|---|---|---|
| 50 to 60% | 233 | 54.3% | 56.2% | +1.9 pts |
| 60 to 70% | 163 | 64.7% | 63.2% | -1.5 pts |
| 70 to 80% | 85 | 74.4% | 70.6% | -3.8 pts |
| 80 to 90% | 41 | 84.0% | 82.9% | -1.1 pts |
| 90 to 100% | 20 | 93.2% | 85.0% | -8.2 pts |
| Sport | Settled | Average stated | Actually won | Brier |
|---|---|---|---|---|
| Darts | 69 | 66.4% | 68.1% | 0.2122 |
| Snooker | 128 | 66.7% | 67.2% | 0.2057 |
| Rugby | 41 | 64.5% | 65.9% | 0.2128 |
| Esports | 76 | 63.4% | 59.2% | 0.2401 |
| Tennis | 228 | 62.5% | 61.4% | 0.2306 |
Brier score measures both accuracy and honesty at once, and lower is better. Predicting 50% on everything scores 0.25, so anything below that is doing real work. It punishes confident wrong answers much harder than cautious ones.
Football's ledger records selections its models flagged, and every one was then rejected by our own published thresholds, so none of them was ever backed. These are the picks we chose NOT to make, which is why they sit apart from the chart above and why their strike rate is not a measure of model performance. For completeness: 102 settled selections, average stated probability 42.6%, actual strike rate 36.3%, Brier score 0.2007. The sample is far too small to conclude anything, and we would say so even if it flattered us.
Calibration is the claim we care most about, because it is the one that makes a probability useful rather than decorative. It is also the one most easily faked by people who quietly drop their losers, which is why every settled prediction we have ever published is in the numbers above, and why the raw data behind them is downloadable as CSV from every record page.
What it does not prove is profit. A perfectly calibrated model can still lose money if the prices are not there, and a badly calibrated one can get lucky for months. Those are separate questions, tracked separately on the record pages. The honest summary today is that the sample is young: bands with fewer than about fifty predictions should be read as indicative rather than settled, and the confidence bars in the chart show you exactly which ones those are.
Generated automatically from the live ledgers on 29 July 2026, so this page cannot drift from the record it describes. Voided and postponed events are excluded; nothing else is. Method: how the models are built and tested.