The policy
Every model has bad days, and every system has bugs. The question that matters is what
happens next. Deleting the affected picks would make the record look better and mean
nothing. So we do the opposite.
- Nothing is deleted, ever. A published prediction stays in the ledger, on the
record page and in the downloadable CSV, whatever it turns out to be worth.
- Faults are written up here. What broke, what it affected, how many records, and
what we changed so it cannot happen again.
- Exclusions are disclosed and counted. If a fault was severe enough that leaving
the affected rows in the headline accuracy would mislead, we may exclude them, and this
page says exactly how many and why. An exclusion you cannot see is a deletion with better
manners.
- The default is to count them anyway. No correction has been excluded from our published statistics. Every affected record is still counted against us.
- A settled row can never change. An automated check compares every settled
prediction against a stored fingerprint each night and alerts if one is altered or
disappears.
Corrections found by readers are welcome and
will be credited. Use the contact form. The fastest way to
make this project better is to catch us being wrong.
Record
9 corrections published ·
0 resulting in any exclusion from headline statistics ·
0 predictions deleted, since day one.
2026-09-15
Annotated, still counted
Snooker likeliest scorelines published for the wrong match length
Snooker
What happened. Our snooker previews take the match length from a table of competition names, and anything the table did not recognise fell back to first to 5 frames. A generic "qualifier" rule also fired before the event-specific ones. The result is that 300 of 363 snooker predictions since July were priced as first to 5 regardless of the event. Where the exchange carried a correct-score market, which states the format exactly, our length was wrong on 16 of the 18 matches we can check: a first-to-10 final was published with "5-2" as its likeliest score, and best-of-7 Home Nations matches were priced as best of 9.
Why it matters. The likeliest scoreline and expected frame count on those cards were arithmetically impossible for the match being played, exactly as happened with the darts World Matchplay in July. The published win probabilities are blended with the market price and are largely unaffected, but on matches without a market price the win probability also depends on the match length and was mis-stated by a few points. The site owner spotted it from the cards: every likeliest score was "5-something".
What we did. Annotated. All affected rows stay in the record with their original scoreline and frame figures; nothing is deleted. On 15 September the twelve Northern Ireland Open qualifier previews published that morning at first to 5 were voided in the ledger (kept, greyed) once the error was found; the five evening matches were re-logged at the correct first to 4 before they started, and the seven afternoon matches were already in play so carry a void with no replacement. Winner picks and their settlement on every other day stand. Scoreline-accuracy figures for snooker before 15 September should be read as unreliable rather than as a measure of the model. From now on a match whose length we cannot verify is published with a winner probability and no scoreline at all.
Fixed. The match length now comes from the exchange's correct-score market wherever one exists, which covers the long-format later rounds where the error was largest. The competition table has been rebuilt with event-family rules ahead of the generic one, Home Nations events at first to 4. Every logged prediction now records where its format came from, and the daily health check raises a warning on any day that used the default. This is the same fix the darts pipeline received on 26 July; it should have been applied to snooker at the same time. Read the detail
2026-07-31
Annotated, still counted
Our indices backtest charged almost no trading cost, so its published Sharpe was overstated
Trading
What happened. The simulator behind the indices and metals book looks the dealing spread up in a table that contains only foreign-exchange pairs. Every index and metal fell through to a default of one FX pip, which is 0.0001 of price. Against real spreads of 0.40 points on the S&P and 22 points on the Nikkei, that is between four thousand and two hundred and twenty thousand times too small, so the backtest paid almost nothing to trade. We had published a Sharpe of 1.19 from it. Going to re-derive it properly, we also found the script that produced that number no longer exists in either repository, so it cannot be reproduced at all.
Why it matters. A backtest that does not pay to trade flatters any strategy that trades often, which is precisely the error that makes a system look better than it is. It also means a number we asked people to judge us on was not supported by anything we could re-run.
What we did. We rebuilt the backtest to mirror what the live bot actually does, charged the real spreads measured from the broker, and published the result even though it is worse: 0.77 across 2011 to 2026, and 1.80 across the held-out era since 2023, at a 29 percent maximum drawdown. Trading costs explain only a small part of the fall from 1.19; the remainder we cannot account for against a script that no longer exists, so we publish the figure we can defend rather than the one we would prefer. The combined three-book figures on the same page still use the old indices input and are labelled as overstated until that blend is re-derived.
Fixed. Indices costs now come from a dedicated model holding real per-instrument spreads, with a double-spread stress column reported alongside, following the pattern the commodities book already used. Overnight financing is still not modelled and is stated as such. The live trading bot is unaffected: it uses the simulator only to decide whether it is currently holding a position, and cost never changes that. Read the detail
2026-07-29
Annotated, still counted
Match-result value bets could still be backed under a threshold our own test had refuted
Football
What happened. The July sweep found no match-result (1X2) value threshold that was profitable on both halves of the held-out data, and we published that on 26 July. The old bar was left active in the code until 29 July, so the system could still have marked such a selection as backed.
Why it matters. Backing a bet under a rule we had already shown does not hold is exactly what our own discipline is meant to prevent, whether or not it costs anything. No match-result value selection cleared that bar between 26 and 29 July, so no published bet or figure is affected.
What we did. Avenue switched off. Those selections stay in the record as tracked, not backed, with their profit and loss still shown.
Fixed. The value avenue is disabled in code with the sweep evidence recorded beside it. The match-result probability avenue (65% at 1.55+), which did pass on both halves, is unchanged. All three unvalidated avenues are re-swept on 1 November. Read the detail
2026-07-29
Annotated, still counted
Accumulator legs now ranked on probability alone
Football
What happened. The big-odds fold applied a small expected-value floor as well as the probability ranking every other tier used. That floor has been removed.
Why it matters. Our accumulator construction rules are published as fixed in advance, so any change to them is recorded here even though it changes no settled bet.
What we did. Recorded. No settled accumulator is altered or removed.
Fixed. Every accumulator tier now selects on model probability inside its fixed odds band, one visible rule per tier. Read the detail
2026-07-26
Annotated, still counted
Placeholder exchange prices recorded as real market prices
Tennis · Darts · Esports · Rugby
What happened. When a betting market has no money in it yet, the exchange shows its minimum price on both sides. Our capture treated that as a real market view. It normalises to a flawless 50/50, so every sanity check we had passed it.
Why it matters. On tennis the fabricated price was not just a benchmark, it was an input to the model, so on those fixtures the model largely echoed a number that meant nothing. 53 of 185 tennis records carry a market price of exactly 50%.
What we did. Annotated. Those rows stay in the record and in the published accuracy figures.
Fixed. Capture now requires a plausible two-sided book AND actual traded volume before storing a price. A daily data-quality check watches for recurrence. Read the detail
2026-07-26
Annotated, still counted
Market price published for the wrong player
Tennis · Darts · Snooker · Rugby · Esports
What happened. On the record pages and downloadable CSVs, the model column showed our favourite's probability while the market column showed the first-named player's. Where the model favoured the second-named player, the two columns described different people.
Why it matters. It made the model look contrarian on matches where it actually agreed with the market. 93 of 325 published rows were affected, the worst by 84 percentage points. The cards were always correct, so cards and pages disagreed.
What we did. Corrected at source and every page regenerated. No prediction changed; only the market figure shown beside it.
Fixed. The publisher now flips the market price to the favourite's side, matching the cards. Verified row by row against the ledgers. Read the detail
2026-07-26
Annotated, still counted
Closing-line figures reported as zero when no closing price existed
Snooker · Esports · Rugby · Darts
What happened. If a closing price was never captured, the code copied the opening price into the closing field. Closing-line value then computed as exactly 0.00% and read as a real measurement rather than a missing one.
Why it matters. Every snooker record was affected. 'Our picks show zero closing-line value' is a meaningful claim about edge; 'we never captured the close' is not, and the two were indistinguishable.
What we did. No prediction affected. The published closing-line coverage figures on the methodology page now state the real per-sport percentages.
Fixed. A missing closing price now stays missing. Coverage is published rather than implied. Read the detail
2026-07-26
Annotated, still counted
Women's World Matchplay priced with the men's final format
Darts
What happened. Our format table matches competitions by name. Betfair calls the event "PDC Womens World Matchplay", which contains "PDC World Matchplay", so the women's first round inherited the men's schedule and was priced as that day's men's final: first to 18 legs, two clear. The women's first round is first to 5, best of 9, with no two-clear rule. Betfair's own correct-score market settles it: its runners run from 5-0 to 5-4.
Why it matters. The published win probabilities were NOT affected, because those come from our blend with the market price rather than from the match length. What was wrong were the likeliest scorelines and the expected number of legs: the card offered "18-10" as the likeliest score in a race to 5, and 29 legs in a best-of-9 match.
What we did. Caught by the site owner on the morning card, roughly two hours after it went out. The four affected rows were voided in the ledger and are still visible there; four fresh rows were logged with the correct format, and a corrected card was reissued to the channel explaining the change. Nothing was deleted.
Fixed. The women's event now has its own entry in the format table, verified against the exchange's correct-score market rather than assumed, and it takes precedence over the men's because the matcher prefers the longest matching name. Rounds we have not verified now raise a visible warning instead of silently falling back to a default. Read the detail
2026-07-17
Annotated, still counted
World Matchplay picks published with the wrong match length, then replaced
Darts
What happened. Eight World Matchplay first-round picks were published on 17 July with the match length set to a default race-to-6 instead of the tournament's actual race-to-10. Match length changes a win probability, so the numbers were wrong: Aspinall v Cullen went out at 74% when the corrected figure was 80%.
Why it matters. The picks were then re-logged with the correct format and the original rows were removed. That was the wrong way to handle it. Deleting and re-logging is exactly the habit this policy exists to replace, and it is why the eight rows are absent from the 17 July record even though a card went out that day.
What we did. Kept as-is and disclosed here. The corrected rows are in the record and settled normally; the superseded originals are not being reinstated because that would put the same eight matches in the record twice. We checked: the model went 7 from 8 on that round, so nothing about this flattered us.
Fixed. From now on a pre-match correction voids the original row visibly and logs a fresh one beside it, so both remain on the record. Match formats are read from the competition rather than defaulted. Read the detail