Insights / Why we stopped publishing a confidence percentage

Why we stopped publishing a confidence percentage

28 August 2026 · 6 min read

Until August, every Algohorse tip carried a percentage: 70% to place. It is gone now. This is what we tested, what we found, and why removing a feature was the right call.

What calibration means

A model is well-calibrated if its stated probabilities match reality: of all the horses it rates at 70%, roughly 70% should actually place. Of the ones it calls 30%, about 30% should. Plot "predicted" against "actual" and a perfectly calibrated model sits on the diagonal. The standard tools are reliability diagrams and summary scores like expected calibration error and the Brier score.

Calibration is necessary, and it is not sufficient

Here is the trap we walked into. Our model is well calibrated across all runners — take every horse in every race and the stated probabilities land close to the observed rates. That is a real property and it is worth having, because it is what makes the ordering trustworthy: the model can tell a 12% chance from a 35% chance.

But the tips we publish are not all runners. They are the handful the model rates highest, and that is a very different population. Being right on average across thousands of horses does not mean being right on the narrow, self-selected slice at the top — and the top is precisely where a model's own optimism concentrates. Selecting on the maximum selects for overestimation.

So we tested it on the tips themselves

The question that actually matters to a member is narrower than "is the model calibrated". It is: among the tips you publish, does a higher stated percentage mean a better chance of placing? That is a discrimination question, and it has a standard measure — the area under the ROC curve. A score of 0.5 means the number carries no information at all; 1.0 means it separates the placers from the rest perfectly.

On our published selections it came out at 0.537, with a 95% range of 0.483 to 0.590. That range includes 0.5. In plain terms: within our own card, the confidence percentage did not tell you which tips were more likely to place. On the top-rated pick specifically the stated probability ran about eight points ahead of the outcome, and in small fields nearer nineteen.

Why we removed it rather than shrinking it

We could have kept the number with a caveat. We did not, for one reason: a percentage on a betting card does not read as a caveat, it reads as a promise. It is the most precise-looking thing on the page, so it anchors the decision — how much to stake, which selection to take if you only take one. A number that looks that precise and carries no information is worse than no number, because it displaces the judgement a member would otherwise make from the price.

What replaced it is the thing we can stand behind: the registered price, the place price alongside it, and how long ours actually stayed available. See the early price and how long it lasts for that measurement, and our record — strike rate and beat-the-close, published with the complete settled ledger behind them. If we find a per-tip confidence measure that survives this test, we will publish the test alongside it.

Start free. Upgrade when the model convinces you.

One selection a day free, plus yesterday's full card graded. The whole card is £14.99 a month.