Quantile Labs Contact
Post

The 94% Problem

Everyone can tell a restaurant with five reviews from one with five thousand, and then loses that intuition the moment the same structure arrives with a percent sign attached. Two evaluations reporting 94% can differ by fourteen points in what they actually support, and the convention for printing results hides the difference from the people deciding.

Contributors
Kayode Adeniyi
Date
Tags
Share
Citation
@misc{quantile-2026-the-94-percent-problem,
  title        = {The 94% Problem},
  author       = {Kayode Adeniyi},
  howpublished = {\url{https://quantilelabs.com/blog/the-94-percent-problem/}},
  year         = {2026},
  month        = {08},
  publisher    = {Quantile Labs},
}

Everyone already understands this problem in a setting where nothing is at stake, because a restaurant with five reviews averaging 4.8 stars and one with five thousand reviews averaging 4.8 stars are plainly not the same restaurant. The intuition is instant and correct, since the second kitchen has been tested often enough that its average has stopped moving, while the first might be excellent or might be four friends on a lucky night.

We use that distinction effortlessly several times a week, then lose it entirely the moment the identical structure appears in a slide deck with a percent sign attached.

The same structure governs how AI systems get reported, with rather more riding on it. An evaluation states that a system was correct 94% of the time and a second reports the identical figure, the first having tested fifty cases and got forty seven right, the second having tested a thousand and got nine hundred and forty. Every published evaluation scheme I have examined prints those two the same way, which hands the reader the restaurant with five reviews and the one with five thousand while calling them equivalent.

The distance between them is where decisions quietly go wrong, because the smaller evaluation supports ninety five percent confidence that the true rate lies between 83.8% and 97.9%, while the larger narrows that range to between 92.4% and 95.3%.

Suppose the contract requires ninety percent, an ordinary sort of bar in a procurement document, in which case the large evaluation clears it on any reading of its own evidence, whereas the small one sits comfortably alongside a system running at 85%, which is to say it is consistent with failing the very requirement it was commissioned to demonstrate.

The reason that intuition deserts people in a meeting has more to do with how numbers are read than with anything mathematical. A percentage is a format that signals precision, precision gets read as confidence, and confidence is what a room under time pressure wants before it can decide and move on.

Nobody in that room is being careless, since they are responding sensibly to a figure that convention has engineered to look finished, the denominator was never part of that format, and a quantity nobody printed cannot be questioned however carefully it is read.

Fields that make consequential decisions from samples settled this a century ago by refusing results in any other shape, and protein structure prediction is the case worth studying, since the results that convinced biologists came out of an assessment run by people with no stake in the entrants, on structures withheld until the predictions were locked in, with stated uncertainty as a condition of entry. The culture around machine learning inherited a narrower purpose, since a leaderboard ranks systems against one another on fixed questions, and for ranking a bare percentage does the job when every entrant answers identical items. The difficulty starts when that number leaves the leaderboard and becomes an absolute statement about a system going into service, carrying a weight nobody built it to bear.

This reaches further than vendors with something to sell, since one serious open source evaluation framework compares a metric against a threshold and reports which side the system landed on, saying nothing whatever about how firmly. A national security institute's standard for evaluation quality, governing what may be submitted to its own testing suite, recommends a binary success metric representing a pass or a fail, and mentions confidence intervals, statistical uncertainty, and reliability nowhere in the document.

The floor is missing directly beneath the people standing closest to the problem, and it has stayed missing because nobody thought to look down.

The practical cost lands on whoever has to decide, because governments and large institutions shape technology markets mainly through what they agree to buy, and that leverage assumes a buyer capable of telling a strong system from a weak one using the evidence in the room.

An official holding 94% is holding an assertion with a percent sign attached, and no diligence recovers from it what was never put in.

Procurement then buys a system that underperforms or declines to buy at all, and both outcomes get filed as evidence that the technology is immature, when the failure sat in the reporting.

A second question hides behind the first, which is who those fifty cases actually were, since a sample supports claims about the population it was drawn from and a deployed system will meet people who were never in that population, some of whom belong to the group where it performs worst, all of whom vanish into an average pooling everybody together. Printing the denominator invites somebody to ask who was counted, whereas a bare percentage closes that conversation before it opens, and the people missing from the count are usually those with the least standing to complain afterwards.

The remedy is dull and available now, beginning with the rule that a result never appears without its denominator and its interval beside it, and where that interval spans a grading boundary an honest report names both grades and declines to choose between them, because selecting either one asserts something the evidence does not contain.

A claim should be capped by the access the evaluator was granted, since testing a system from outside cannot support conclusions requiring a view of its inside, and the package should be sealed so somebody else can recompute every figure from the underlying observations, offline, years later, without extending trust to whoever produced it.

The statistics involved are a century old, the particular interval quoted above was published in 1927, and nothing here waits on a research breakthrough or a larger budget.

What has happened is that a technology now deciding who receives credit, who is seen first in a hospital, and who keeps a benefit has arrived inside an evaluation culture that never had to grow up, and the figures reaching the people who decide would not survive a first year methods seminar.

The problem is embarrassing for the same reason it is fixable, which is that everybody already owns the intuition and uses it whenever they book dinner.