unspurious.

The inference illusions · The Law of Small Numbers

The best school and the worst are both tiny.

Rank anything — schools, towns, hospitals — and the top and bottom of the table fill up with the smallest groups. Not because small is better, or worse, but because a small sample swings further from the truth. The record is a story about size.

A county’s schools, ranked by top grades Each dot is one school: how big it is, and the share of its pupils who got top marks.
every school all sizes

Record high Record low Ordinary school The real rate: 30%

Top of the table
Bottom of the table
30%
The real rate
the same in every school

Fig. 1 — The funnel of luck. Every school here has the same true ability: 30% of pupils get top marks, set by a coin weighted 30/70. But small schools (left) ride their luck — a few bright or unlucky pupils swing the whole percentage — while big schools (right) average theirs out and hug 30%. So the table’s record highs and record lows are all crowded at the small end. Raise the size cutoff and the drama drains away.

The short answer

What is the law of small numbers?

The law of small numbers is the mistaken belief that a small sample will look just like the population it comes from. It doesn’t: small samples vary far more than large ones, so their averages swing to extremes. The name is a joke on the genuine law of large numbers, coined by Amos Tversky and Daniel Kahneman in 1971 for the way people wrongly expect small samples to be reliable.

The question that saves you

How big is the sample behind this record?

Whenever something tops a ranking — the safest town, the best-performing school, the ward with the lowest cancer rate — the first question is not “what are they doing right?” but “how many cases is that built on?” A record set by a small sample is the single most likely thing you will ever see, because small samples are where luck runs unchecked. The extremes of almost any list are a roll-call of its smallest members.

AskIs this a real difference, or just a small sample free to wander?

01 · The map that fooled everyone

Why the healthiest counties are also the sickest

Here is a genuine finding that has misled careful people for decades. Map the rate of kidney cancer across the counties of the United States and colour the lowest ones green. They turn out to be mostly rural, sparsely populated and in the Midwest, the South and the West. It is an easy story to tell: clean country living, fresh air, no city stress. Now colour the highest counties red. They are also mostly rural, sparsely populated and in the Midwest, the South and the West. The same kind of county sits at both ends.

No lifestyle explains that, because both patterns have the same cause, and it has nothing to do with kidneys. A rural county might have only a few thousand people. In a small population, one or two cases — or none — swings the rate enormously. Toss a coin four times and two heads is unremarkable; four heads happens one time in sixteen. Toss it four thousand times and you will never see all heads. The rate isn’t high or low because of anything the county does. It is high or low because the county is small, and small numbers are free to wander.

This is the psychologists’ “belief in the law of small numbers” — Tversky and Kahneman’s 1971 name for our habit of expecting a handful of cases to be as trustworthy as a multitude.

02 · The most dangerous equation

Why averages of a few things swing so hard

There is one short formula underneath all of this. The spread of an average — its standard error — is the spread of the individual things divided by the square root of how many you averaged. The statistician Howard Wainer called it “the most dangerous equation”, not because it is hard, but because not knowing it has cost fortunes and misdirected policy for three centuries.

The square root is the whole story. Average four pupils and the result scatters twice as widely as an average of sixteen; average four hundred and it scatters ten times less than the four. Halve the sample and you don’t double the noise, but you do inflate it — and shrink the sample to a handful and the noise can drown the signal completely. That is why the funnel above has the shape it does: a wide, wild mouth on the left where schools are small, narrowing to a tight neck on the right where they are large. The true rate is a flat line straight through the middle the whole way. Nothing about ability changes across the chart — only how much room luck has to move.

03 · The billion-dollar mistake

How small samples cost the schools $1.7 billion

This is not a museum piece. In the late 1990s the Bill & Melinda Gates Foundation, among the best-resourced philanthropies on earth, noticed that small schools were strikingly over-represented among the country’s highest achievers. The conclusion seemed obvious and humane: big schools were failing children, so break them up. By 2005 the foundation and others had poured some $1.7 billion into creating small schools.

The trouble was the other end of the list. Small schools were also over-represented among the worst performers — for exactly the reason the funnel shows. With few pupils, a single strong or weak cohort moves the whole average, so small schools crowd both extremes at once. The high performers weren’t small because small is good; they were the lucky tail of a distribution whose unlucky tail was small too. In 2005 the foundation quietly shifted away from the small-schools push. A more reliable predictor of a school’s scores, it turned out, was simply how many pupils it had — the one variable nobody was supposed to be studying.

Chase the top of a noisy ranking and you are really selecting for small samples. Fund them, copy them, promote them — and next year a fresh set of small samples takes their place.

04 · How not to be fooled

Reading a record safely

Find the denominator. Behind every rate is a fraction; the number on the bottom is the one that matters here. “Cancer rate 3× the national average” means one thing in a city of a million and nothing at all in a hamlet of eight hundred. If a report gives you the rate but hides the sample, it has hidden the only thing you need.

Distrust the extremes of any list. The top and bottom of a ranking are where small samples gather. The safest, most reliable-looking members of a league table are usually in the crowded middle, built on enough cases to have nowhere to hide.

Expect the record to fade. A number that is extreme mostly because it is small will drift back toward the average the moment you measure it again — the trap of regression to the mean. This is why last year’s miracle school so often looks ordinary the year after.

Weight by size before you rank. A funnel plot — the picture above — is the honest way to show rates against sample size, so a small group’s wild number can’t masquerade as a real difference. Compare like with like, and give a big sample the greater say it has earned.

Continue the field guide

More ways noise dresses up as signal