unspurious.

The inference illusions · The prosecutor’s fallacy

The chance of the evidence is not the chance of innocence.

Tell a jury that only one innocent person in a million would match the evidence, and they hear a one-in-a-million chance the defendant is innocent. Those are different numbers — and the gap between them has sent innocent people to prison.

Everyone the evidence fits A trace at the scene matches the defendant. How many other people does it also match — and how much does that change things?

Someone who also matches The defendant in the dock

“Chance he’s innocent”
what the court is told
Others who also match
innocent, in the pool
Chance it’s actually him
on this match alone

Fig. 1 — One match, a crowd of suspects. The expert’s number is honest: it really is the chance an innocent person matches. But guilt depends on a second number the expert never gives — how many people could have done it. Spread the same rare match across a city and it fits a roomful of innocents; concentrate it in a village and it points almost squarely at one. (We assume here that the real culprit is somewhere in the pool and does match.)

The short answer

What is the prosecutor’s fallacy?

The prosecutor’s fallacy is the mistake of confusing two different probabilities: the chance that an innocent person would match the evidence, and the chance that a matching person is innocent. An expert might say only one innocent person in a million would match a DNA sample. That is not the same as a one-in-a-million chance the defendant is innocent, because in a large population many innocent people can match by chance. Swapping the two is called transposing the conditional.

The question that saves you

Did they give you the chance of the evidence, or the chance of innocence?

An expert can tell you how often a piece of evidence turns up by chance among innocent people — “one in a million”. That is not how likely the defendant is to be innocent. To get there you also need two things the rare number quietly skips: how many people could have done it (a rare match in a big crowd still fits many of them), and how rare the innocent explanation is compared with the guilty one. Leave those out and a frightening-sounding probability slips into the seat reserved for a verdict.

AskThe chance of what, given what — and out of how many possible people?

01 · The swap

Two questions that sound identical — and aren’t

Most professional basketball players are tall. That does not make most tall people basketball players — there are millions of tall people and only a few hundred professionals. Flip the order of a “most A are B” statement and you can get a completely different, usually wrong, answer. Statisticians call it transposing the conditional; everyone else has just felt the trick land.

A courtroom runs the same swap with probabilities. The honest claim an expert can make is:

the chance an innocent person would match  =  1 in a million

What the jury hears, and what the prosecutor often implies, is the reversed version:

the chance a matching person is innocent  =  1 in a million

They are not the same statement, and the gap between them can be the difference between a handful and a city. Whether the second is even close to true depends entirely on something the first never mentions: how many people there were to begin with. In a population of millions, a “one in a million” match is something you should expect several innocent people to have.

It is the same shape of error as the base-rate fallacy in a hospital: “the test is 99% accurate” is not “your positive result is 99% likely to be real”. Rarity of the evidence is only half the sum.

02 · What it cost

Sally Clark, and a number that should never have been said

In 1999 Sally Clark, an English solicitor, was convicted of murdering her two baby sons, who had died suddenly some weeks apart. A central plank of the case was a single statistic. A paediatrician told the jury that the chance of two cot deaths (sudden infant deaths) in a family like hers was 1 in 73 million — a number arrived at by taking a one-in-8,500 chance for a single cot death and simply multiplying it by itself.

It was wrong twice over, and each error is one this site is about.

The first error: multiplying as if the deaths were independent. Squaring 1 in 8,500 only works if the two deaths had nothing to do with each other — like two separate coin flips. But siblings share genes, a home and an environment, so a family that suffers one cot death is, sadly, more likely than average to suffer another. Once you allow for that, the real chance of a second is far higher and the “1 in 73 million” collapses.

The second error: the prosecutor’s fallacy itself. Even if two natural deaths really were that rare, that is still not the chance Sally Clark was innocent. To reach that, you have to ask how likely the alternative is — a mother murdering two of her own babies — and that, mercifully, is also vanishingly rare. The right comparison is between two rare tragedies, and when statisticians later did the comparison properly, double natural death came out as the more likely of the two. The “1 in 73 million” answered a question that should never have been asked.

The Royal Statistical Society took the unusual step of writing publicly about the misuse of statistics in the case. Clark’s conviction was quashed in 2003, after she had spent more than three years in prison. She never recovered, and died in 2007. The number that convicted her was not a lie; it was a true answer to the wrong question.

03 · Rare is only half the question

Why both rare things have to be weighed

The honest question is never “how unlikely is this evidence if he’s innocent?” on its own. It is comparative: how much more likely is the evidence if he’s guilty than if he’s innocent? — and that ratio updates a starting point set by how many people could have done it.

Put concretely: a one-in-a-million match multiplies the odds of guilt by about a million. That sounds decisive until you remember where the odds started. If the trace could have come from any of eight million people in a city, you begin at roughly one in eight million and the match carries you to about one in eight — powerful, but a long way from certainty. If other evidence has already narrowed the field to a village of a few thousand, the same match lands you at near-certainty. The match is exactly as strong as the pool is small. You watched that in the figure: widen the pool and the defendant disappears into a crowd of equal matches; shrink it and he stands alone.

This is also why trawling a giant DNA database for a match needs special care. Search a database of millions and you should expect to turn up a chance match or two, exactly as testing twenty hypotheses hands you a “significant” one — the Texas sharpshooter with a forensic lab. A cold-hit match found by searching is weaker evidence than the same match on a suspect singled out for other reasons, and the maths has to say so.

04 · Neither fallacy

How to read a “one in a million” safely

There is an opposite mistake, and the defence makes it: “thousands of people share this profile, so the match means nothing.” That is the defence attorney’s fallacy, and it is just as wrong. A match that fits one person in a million is genuinely powerful — it cuts the field from millions to a handful. The truth sits between the two fallacies: the match is strong evidence, but it is one input, to be combined with everything else and weighed against the size of the suspect pool. It updates the odds; it does not deliver the verdict.

So when a single probability is offered as proof, take it apart:

Ask “the chance of what, given what?” Make them say whether it is the chance of the evidence assuming innocence, or the chance of innocence given the evidence. Only the second bears on guilt, and it is almost never the number quoted.

Ask how big the pool of possible suspects is. The same match means near-certainty in a village and very little in a nation. No pool size, no conclusion.

Demand the comparison. How likely is the evidence if guilty versus if innocent — and how likely was guilt before this evidence? A probability with nothing to compare it against is theatre.

Be suspicious of multiplied probabilities. “One in a thousand, times one in a thousand, is one in a million” only holds if the events are independent. Shared causes — genes, environment, a common source — break the multiplication, as they did for Sally Clark.

It has a long rap sheet beyond that one case: People v. Collins (California, 1968), where a couple was convicted on multiplied eyewitness odds and freed on appeal; the wrongful conviction of the Dutch nurse Lucia de Berk, cleared in 2010 after a “1 in 342 million” coincidence fell apart. The names change; the swapped conditional does not.

Continue the field guide

More ways probability gets read backwards