unspurious.

The selection illusions · Self-selection bias

Your poll only heard from people who wanted to answer.

Online polls, star ratings and “tell us what you think” surveys collect the loud, the angry and the delighted — the ones who opted in. That group is reliably different from everyone who stayed silent, so the numbers describe the volunteers, not the world.

1,000 customers — and who left a review Everyone had an experience. Only some wrote it down, and who writes is not who bought.

Everyone who bought Those who reviewed True average Review average

The star average shown
what you see on the page
Customers who replied
the self-selected few
The true average
everyone’s experience

Fig. 1 — The reviews are real; the sample is not. Every customer truly had the experience in the faint bars. But the solid bars are who bothered to write it up — and when the quiet middle stays home, or one camp shouts louder, the review average drifts away from the true average nobody mistyped a single number to produce.

The short answer

What is self-selection bias?

Self-selection bias happens when the people in a sample choose for themselves whether to take part. Because the ones who opt in — to a poll, a review, a survey — tend to feel more strongly or differently than the ones who stay silent, the sample no longer represents the whole group. The result describes the volunteers, not the population.

The question that saves you

Who didn’t answer — and how are they different?

A poll, a review section or a survey only ever contains the people who chose to fill it in. If those people feel differently from the ones who didn’t — and they almost always do, because caring enough to reply is itself an opinion — then the result is bent before you read a single number. The fix is never a bigger pile of the same volunteers; it is a sample you chose, with as few people missing as possible.

AskDid I choose this sample, or did it choose itself — and who is silently missing?

01 · The mechanism

A sample that picks itself is already bent

A fair sample is one you draw — ideally at random — so that everyone in the population has the same chance of being in it. Self-selection turns that on its head: the sample decides for itself who joins. Whoever clicks the poll, writes the review, returns the card or volunteers for the study puts themselves in the data, and everyone who shrugs and moves on quietly leaves it.

That would be harmless if the people who reply were just a smaller copy of everyone else. They never are. Replying takes a flicker of motivation, and motivation is exactly the thing under measurement. The people moved enough to act — thrilled, furious, passionate, aggrieved — are by definition unlike the larger middle who felt the situation didn’t warrant the bother. So the gap between responders and non-responders isn’t random noise; it points in a direction, and it bends the answer that way.

It is a cousin of survivorship bias: there, the data you see is whatever survived a filter; here, it is whoever volunteered. Both times, the missing people are the ones with the most to tell you.

02 · Two and a half million wrong answers

The poll that was huge, confident, and badly wrong

In 1936 the Literary Digest, a respected American magazine with a perfect record of calling elections, ran the biggest poll the world had seen. It mailed about ten million straw-vote ballots and got back roughly 2.4 million. On that mountain of data it announced that Alf Landon would beat Franklin Roosevelt by 57% to 43%.

Roosevelt won 61% to 37%, carrying 46 of the 48 states — one of the largest landslides in American history. The Digest, humiliated, folded within two years.

It failed twice, and both failures are self-selection. First, the list: the ballots went to people from telephone directories, car-registration records and the magazine’s own subscribers. In the depths of the Depression, owning a phone, a car or a magazine subscription marked you as comfortably off — and the well-off leaned to Landon. Second, the returns: only about a quarter of the ballots came back, and the people who felt strongly enough to post them — disproportionately Landon’s motivated supporters — were not the quarter who didn’t.

Meanwhile a young pollster named George Gallup surveyed only about 50,000 people — but chose them to mirror the country — and called the result correctly. Here is the lesson with the paint stripped off: 2.4 million self-selected answers lost to 50,000 representative ones. A bigger biased sample is not closer to the truth; it is a more confident measurement of the wrong thing.

03 · Where it bites

The volunteers are everywhere now

The Digest mailed ballots; today the ballots mail themselves, constantly, and almost every “what people think” number you meet is self-selected.

Online polls. A poll on a website or social feed measures who follows that account and cared enough to tap — never the public. A landslide in a Twitter/X poll is a fact about that follower base and nothing more.

Star ratings and reviews. Most buyers never review. The ones who do skew to the delighted and the outraged, which is why ratings so often pile up at five stars and one star with a hollow middle — the shape you can dial in above. The headline average can sit a full star away from the typical customer’s experience, in either direction.

Feedback and satisfaction surveys. “How did we do?” forms, app-store prompts and net-promoter scores all hear loudest from the extremes. A flood of complaints after a change may be a vocal few, not a verdict — and a wall of praise may be your fans, not your users.

Call-in and write-in everything. Talk-radio votes, petitions, comment sections, “email us your view” — all of them are survivorship’s chatty sibling: a record of who was moved to speak, mistaken for a record of what people think.

Even research. Studies that recruit volunteers, or whose participants can drop out, inherit the same crack: those who sign up and stay are not those who don’t.

04 · How not to be fooled

Reading a number that picked itself

Ask the response rate. If 5% replied, the other 95% are the story, and you have no idea where they sit. A low response rate is a flashing light, not a footnote.

Ask who is missing, and why. Name the people who couldn’t or wouldn’t answer — no phone, no time, no strong feeling, no account — and ask whether they’d have answered differently. If yes, the result leans the other way.

Distrust a sample that chose itself, however big. “Over a million responses” is not a strength when the million chose to be there. Size buys precision, not representativeness; it cannot rescue a bent sample, as the Digest learned with millions to spare.

Look for the silent middle. When ratings or opinions split into two loud camps with an empty centre, suspect that the moderate majority simply never showed up — and that the “average” describes almost nobody.

Prefer a sample someone chose. A random, representative sample with a high response rate — the dull, expensive kind — is the only one that speaks for the whole. When you can’t have it, say what you’ve actually got: the view of the volunteers.

Continue the field guide

More ways the data you see is filtered