StatGardenREF. DESK
Calculators/Blog/A 99% Accurate Test Can Be Wrong Five Times Out of Six
Blog

A 99% Accurate Test Can Be Wrong Five Times Out of Six

The accuracy of a test tells you almost nothing on its own. What matters is how many people being tested actually have the thing.

Published 26 September 2026

Here is a result that reliably surprises people, including professionals who ought to know better.

A test catches 99 per cent of genuine cases and wrongly flags only 5 per cent of healthy people. The condition affects 1 per cent of the population. You test positive. The probability you actually have it is 16.7 per cent.

Not 99 per cent. Not 95. About one in six.

Count the people, not the percentages

Percentages hide what is happening. Counts make it obvious.

Test 10,000 people. About 100 have the condition, and the test finds 99 of them. The other 9,900 do not have it, and the test wrongly flags 5 per cent of them, which is 495 people.

So 594 people test positive, and only 99 of them are genuine. Ninety-nine out of 594 is 16.7 per cent.

The false positives outnumber the true ones five to one, and not because the test is bad. They outnumber them because there are 99 times more healthy people available to be wrongly flagged.

Move the base rate and everything changes

The test is a fixed thing. Its sensitivity and specificity are properties of the test and do not change. What changes the answer is who you point it at.

At a 0.1 per cent base rate, the same test gives a posterior of 1.94 per cent. At 5 per cent it gives 51 per cent. At 50 per cent it gives 95.2 per cent.

Nothing about the test differed between those. The Bayes theorem calculator takes the prior as its first input for exactly this reason: without it, the accuracy figures cannot be turned into an answer at all.

This is why screening and diagnosis are different activities

Screening applies a test to a general population where the condition is rare, so the base rate is low and most positives are false. Diagnostic testing applies it to people with symptoms, where the base rate is far higher and a positive means much more.

Same test, same accuracy, genuinely different information. It is why a positive screening result is a reason for a confirmatory test rather than a diagnosis, and why expanding screening to lower-risk groups produces more false alarms rather than more detection.

Accuracy is the wrong summary

A single accuracy percentage is close to useless for a rare condition. A test that simply declares everyone negative is 99 per cent accurate when the condition affects 1 per cent of people, while finding nobody at all.

Sensitivity and specificity separate the two failure modes, and the diagnostic test metrics calculator reports both alongside the predictive values. The predictive values are the ones that answer the patient's question, and they are the ones that move with the base rate.

It generalises well beyond medicine

Any rare-event detector has this shape: fraud flags, spam filters, security alerts, quality inspections. When the thing you are looking for is rare, most of what the detector reports will be false, however good the detector is.

That is not an argument against detecting. It is an argument for treating a flag as a reason to look more closely rather than as a conclusion, and for designing the follow-up step deliberately rather than treating the alert as the answer.

For conditional probability generally, see the conditional probability calculator. For a different way intuition fails on probability, see the piece on gambler's fallacy.

Advertisement
Advertisement