Media & Information
Polls and Statistics
Reading polls and numbers without being misled.
126 min read · 27,800 words
“It is the mark of a truly intelligent person to be moved by statistics.”
Foreword
In 1936, the Literary Digest, a popular American magazine, conducted what was then the most ambitious poll in history. The magazine sent more than ten million ballots through the mail asking Americans whether they would vote to re-elect President Franklin Roosevelt or to elect his Republican challenger, Alf Landon. More than two million ballots came back. The magazine, processing this enormous sample, declared with confidence that Landon would win in a landslide — carrying 57 percent of the vote. Roosevelt won 61 percent and forty-six of forty-eight states. The Literary Digest folded within two years.
That same year, a young pollster named George Gallup, working with a sample of fewer than fifty thousand respondents — a fraction of the Digest’s sample — correctly predicted Roosevelt’s victory. Gallup understood something the Digest had missed: a sample is only as good as its representativeness, not its size. The Digest had drawn its sample from telephone directories and automobile registration lists, which in 1936 systematically excluded the poorer Americans most likely to vote for Roosevelt. The two million returned ballots were a vast sample of the wrong people.
This guide is, in some ways, an extended meditation on the lesson of 1936. Polls and statistics are powerful tools for understanding what is happening in a country too large for any of us to see directly. They are also, when done badly, powerful tools for misleading. The difference between good and bad polling is often invisible to the casual reader; the numbers look the same. The work of citizenship in a country governed by polls and statistics is partly the work of learning to tell the difference.
This is a practical guide for ordinary citizens, not a technical manual for survey researchers. The aim is to give the reader enough understanding of how polls and statistics work to read them with appropriate skepticism, recognize the most common errors and manipulations, and know what additional questions to ask before drawing conclusions. The chapters move from the basic mechanics of polling, through the specific challenges of election polls and other survey types, into the broader question of how citizens should think about statistical claims in public life. Each chapter is short enough to read in a single sitting and includes pointers to deeper sources for readers who want to go further.
A few framing principles. First, this guide is non-partisan. Statistical literacy is not a partisan virtue; numbers can be misused to support any political position, and a clear-eyed citizen needs to recognize the misuse regardless of the political direction. Second, the guide is realistic about both the power and the limits of polling. Polls are real instruments for understanding public opinion; they are also imperfect instruments that fail in particular ways. Honest engagement with polling means acknowledging both. Third, the guide proceeds from the conviction that statistical literacy is part of civic competence. A citizen who cannot read a poll — who cannot tell whether the result is meaningful, whether the question was loaded, whether the sample was reasonable — is a citizen at the mercy of whoever is using the numbers.
The American philosopher John Dewey wrote that democracy depends on “the creation of a true public.” A true public is one capable of forming its own judgment from the available evidence. In an information environment saturated with poll numbers, statistical claims, and confidently asserted percentages, the formation of independent judgment requires understanding how the numbers are produced and what they actually mean. This guide offers what tools it can toward that work.
PART ONE
What Polls Are and How They Work
The basic mechanics: what a poll is, how sampling works, and what the margin of error does and does not tell us
CHAPTER 1
What a Poll Actually Is
A poll is a structured attempt to learn what a large group of people thinks by asking a smaller, carefully selected group of people. The question — “What percentage of Americans approve of the president’s handling of the economy?” — is impossible to answer by asking every American. There are too many of them, they live in too many places, and the cost of a true census of opinion would be prohibitive even for the federal government. The alternative, developed and refined over the twentieth century, is to ask a sample of Americans and infer from their responses the views of the larger population. When the sample is well-constructed and the questions are well-asked, the inference is reasonably accurate. When either fails, the inference fails with it.
The basic structure
Every poll has the same basic elements, however much the technical details vary. A polling organization defines a target population (“American adults,” “registered voters in Michigan,” “likely Republican primary voters”). It develops a sampling frame, which is a practical method for reaching members of that population. It selects a sample from the frame using some procedure (random selection, panel-based selection, opt-in selection). It develops a questionnaire — the actual questions to be asked, in their actual order, with their actual wording. It administers the questionnaire to the sample. It tabulates the responses, applies adjustments to correct for known imbalances in who responded, and reports the results, typically with a margin of error and methodological notes.
Each step in this process is a potential source of error. The target population may be defined in a way that produces misleading results (e.g., “registered voters” as a proxy for “actual voters” in an election whose turnout will be uncertain). The sampling frame may exclude or under-represent significant groups (the Literary Digest’s automobile-and-telephone frame in 1936). The selection procedure may produce biased samples (an opt-in online panel that attracts particular kinds of respondents). The questionnaire may push respondents toward particular answers (a question worded to suggest the desired response). The administration may produce a low response rate that distorts the sample (a phone poll where 95 percent of those called refuse to answer). The adjustments may introduce or fail to correct for biases (a weighting scheme that does not account for an important variable). The reported margin of error may convey more confidence than the data actually warrants. Reading polls carefully means thinking about all of these stages, not just the headline number.
What polls measure: opinions, behaviors, predictions
Polls measure several different kinds of things, and the differences matter for how to interpret them:
- Opinions and attitudes. “Do you approve or disapprove of the president’s handling of the economy?” “Do you support or oppose the proposed law?” Opinion polls aim to capture what people think about something at a given moment. Subject to the question-wording problems discussed in Chapter 4, these can be reasonably accurate as snapshots of the moment, though opinions can change quickly and a single poll is just one moment.
- Self-reported behaviors. “Did you vote in the last election?” “How often do you attend religious services?” “Do you drink alcohol?” Behavioral self-reports are typically less accurate than respondents believe themselves to be; people overstate behaviors they consider socially desirable (voting, exercise, religious attendance) and understate those they consider undesirable (drug use, infidelity, racial prejudice). Some of this can be corrected with careful question design and validation against external records, but the underlying problem is real.
- Predictions. “Who will you vote for in the upcoming election?” “Which candidate is more likely to handle the economy well?” Predictive questions ask people to project forward in time, which they do with varying accuracy. The classic case is election polling: people’s stated voting intentions before an election are reasonably predictive of their actual votes, but with substantial error, and the error has historically been larger in close elections than in landslides.
- Knowledge tests. “Can you name the three branches of the federal government?” Knowledge questions measure what respondents know rather than what they think. These are useful for understanding what citizens are working with as they form opinions; results have been consistently sobering.
- Hypothetical scenarios. “If the election were held today, would you vote for X or Y?” “Would you support the law if it included Z provision?” Hypothetical questions are essentially predictions about counterfactual situations and tend to be the least reliable. Respondents do not always know how they would actually behave in a hypothetical situation; the gap between stated and actual behavior can be substantial.
What a single poll is and is not
A single poll is a snapshot of one sample of one population at one moment in time, with one specific question wording, conducted with one specific methodology, by one specific organization. It is informative but limited. Reasonable readers of polls treat any single poll as one data point among many, look for patterns across multiple polls (averages, trends), and read each poll within the context of its specific methodology and source. The headline of a single poll is rarely the right place to draw a strong conclusion.
This is one reason polling aggregators — sites like FiveThirtyEight, RealClearPolitics, Silver Bulletin, and others — emerged as important resources. By averaging many polls, with adjustments for the historical accuracy and house effects of different pollsters, aggregators can provide more reliable estimates than any single poll. They have their own limitations (Chapter 9 discusses these), but they are typically more useful as a guide to the actual state of opinion than any individual poll.
The vocabulary
Reading polls carefully requires familiarity with some basic vocabulary. The Glossary at the back of this guide defines these in detail; here are the essentials:
- Population. The full group whose views the poll wants to estimate. Examples: American adults, Texas registered voters, Republican primary voters in New Hampshire.
- Sample. The subset of the population that is actually surveyed.
- Sample size. The number of completed responses, typically between 800 and 3,000 for national polls. The relationship between sample size and accuracy is non-linear; doubling the sample only modestly improves precision.
- Margin of error. A statistical estimate of how much the sample’s results might differ from the full population’s views due to random sampling variation. Discussed in detail in Chapter 3.
- Confidence level. Almost always 95 percent in published polls, meaning that if the same poll were repeated many times, 95 percent of the resulting margins of error would contain the true population value.
- Response rate. The percentage of those contacted who completed the survey. Discussed in Chapter 5; declining response rates are one of the central challenges of contemporary polling.
- Weighting. Statistical adjustment of the raw responses to make the sample better match the demographic composition of the target population. Discussed in Chapter 6.
- Likely-voter model. For election polls, a method for estimating which respondents will actually turn out to vote, since polls of all registered voters typically over-represent groups less likely to vote. Discussed in Chapter 7.
- House effect. The tendency of a particular polling organization to consistently produce results that lean in one direction relative to other pollsters.
The bottom line
A poll is a structured attempt to learn what a large group thinks by asking a smaller, carefully selected group. Each stage — defining the population, drawing the sample, asking the questions, processing the responses — is a potential source of error. A single poll is a snapshot subject to many forms of uncertainty; multiple polls and aggregators provide more reliable estimates than any single poll. Reading polls carefully means understanding the basic vocabulary, knowing what specific kinds of measurement they are doing, and treating any individual poll as one data point rather than a final answer.
What to read or watch next
- Pew Research Center, “Methods 101” series (pewresearch.org/methods). Free, accessible explainers from one of the most respected polling organizations. The video series and short articles are an excellent introduction.
- American Association for Public Opinion Research (aapor.org). The professional association of public opinion researchers. The site offers transparency standards, post-election polling reviews, and guidance for journalists and the public.
- Robert M. Groves, Floyd J. Fowler Jr., et al., Survey Methodology (Wiley, 2009; updated editions). The standard textbook for survey research; technical but the framework is invaluable.
- Andrew Gelman et al., Regression and Other Stories (Cambridge University Press, 2020). For readers wanting the statistical foundations from a leading practitioner.
CHAPTER 2
Sampling: How a Few Stand for the Many
The central insight of modern polling is counterintuitive: a small, carefully selected sample can tell you, with reasonable accuracy, what a much larger population thinks. A poll of fifteen hundred Americans, properly drawn, can estimate the views of the country’s 260 million adults to within a few percentage points. This is not magic; it is mathematics, specifically the mathematics of probability sampling. But it depends critically on the word “properly.” An improperly drawn sample of the same fifteen hundred can be wildly misleading, regardless of how carefully the analysis is done after the fact. This chapter is about what proper sampling actually means, why it matters, and why it has become harder to do as the twentieth-century methods have eroded.
The fundamental principle: every member of the population must have a known chance of being selected
In a probability sample, every member of the target population has some known, non-zero chance of being included in the sample. This is what makes statistical inference possible. If you know the chance each person had of being selected, you can mathematically reason backward from the sample’s composition to the population’s composition. If you do not know those chances — if some kinds of people had a higher chance of being selected than others, in ways you cannot quantify — the inference breaks down.
The Literary Digest’s 1936 disaster illustrates the principle in failure. By drawing its sample from telephone directories and car registrations, the Digest gave higher chances of selection to people who could afford telephones and cars in the depths of the Depression — disproportionately wealthier Americans, disproportionately Republican. Poorer Americans, who could not afford these things, had near-zero chances of being selected. The resulting sample of millions was systematically skewed; no amount of post-survey analysis could correct for the basic flaw in the sampling frame.
Random sampling and its alternatives
The gold standard for survey research has long been random-digit dialing or random selection from voter registration lists or address-based samples. In each case, the goal is the same: to give every member of the target population (or a close approximation of it) a roughly equal chance of being contacted. When this works, the resulting sample is representative of the population, and the statistics work out as the textbooks describe.
In recent years, the practical difficulties of probability sampling have grown. Random-digit dialing of landlines now reaches a shrinking and unrepresentative slice of the population (older, more rural, more often non-mobile-only). Cell-phone polling is harder; people are less likely to answer unknown numbers, and federal regulations restrict automated dialing. Address-based sampling with mail follow-up is expensive and slow. Online panels recruited from probability samples can preserve some of the statistical properties of probability sampling, but rely on the panelists’ continued willingness to participate.
As a result, many polls today use opt-in online panels recruited from sources like online ads, marketing databases, or volunteer pools. These are not probability samples in the technical sense; the chance that a given American has of being contacted by such a poll is unknown and varies dramatically by their internet behavior, marketing-database presence, and willingness to take online surveys. Pollsters using these panels apply elaborate statistical adjustments to make the resulting samples look like the population on observable characteristics, but the adjustments cannot fully correct for the unobservable differences between people who join opt-in panels and those who do not.
How big a sample do you need?
The relationship between sample size and accuracy is one of the most counterintuitive features of polling. Bigger samples are not proportionally better. Doubling a sample from 500 to 1,000 reduces the margin of error meaningfully. Doubling again from 1,000 to 2,000 reduces it less. Doubling again to 4,000 reduces it less still. The mathematics produces diminishing returns; beyond about 1,500 to 2,000 respondents for a national poll, additional sample size yields modest improvements at substantial cost.
The much larger source of error in modern polling is not sample size but sample bias — the systematic differences between who responds and who does not. A sample of 50,000 with a strong response bias is less accurate than a sample of 1,500 without one. The Literary Digest had two million respondents and was off by 19 percentage points; Gallup had fewer than 50,000 and was right. The lesson has been relearned in every generation since.
This is one reason aggregating polls is more useful than just looking at the largest one. Multiple smaller polls, each drawn somewhat differently, often combine to produce a more accurate picture than any single large poll, because the biases of any one methodology can be partially canceled by the biases of others.
Subgroup analysis: a special caution
Polls often report results for subgroups: “women voters,” “voters under 30,” “white evangelical Protestants,” “Hispanic voters in Florida.” These subgroup results are typically much less reliable than the overall results, for a simple reason: the subgroup is a smaller sample, with a larger margin of error. A poll with 1,500 respondents and a margin of error of ±2.5 percentage points might have only 150 Hispanic respondents, with a subgroup margin of error of ±8 percentage points or worse. Reporters and analysts often slice polls into subgroups without noting that the subgroup results are essentially noise.
The problem compounds when subgroups are reported in the news as showing dramatic shifts. A poll showing that “Hispanic voters under 30 in the Sun Belt” shifted by ten points might be entirely consistent with no actual shift, given the tiny subgroup sample. The same caution applies to changes in subgroup support over time: noise looks like signal when the subgroups are small.
What probability sampling cannot fix
Even with a perfectly executed probability sample, polls measure what people are willing to say to a stranger asking questions. People may not want to admit views they consider stigmatized; people may not have settled views on the question being asked; people may answer differently depending on the wording, the order, the context. None of these problems are solved by larger samples or better sampling techniques; they are intrinsic to the method of asking people what they think. Chapter 4 discusses the question-wording dimension; Chapters 5 and 6 discuss the response and adjustment problems. The honest takeaway is that polls are a useful but limited instrument for understanding public opinion, and the limits are real.
The bottom line
Sampling is the engine of polling; without representative samples, the entire enterprise breaks down. Every member of the target population must have a knowable chance of selection. Probability samples are the gold standard but have grown harder to execute as response rates have collapsed. Larger samples are not proportionally better; the relationship is non-linear and modest beyond about 1,500. Subgroup analyses are typically much less reliable than headlines suggest, because subgroups are smaller and noisier. Even perfect samples cannot fix the underlying problems of measuring what people will say to a stranger.
What to read or watch next
- Andrew Gelman, John B. Carlin, et al., Bayesian Data Analysis (CRC Press, 2013, 3rd ed.). The standard for survey weighting and post-stratification in modern polling, written by some of the field’s leaders.
- Sharon Lohr, Sampling: Design and Analysis (Chapman and Hall, 2019). Comprehensive textbook on sampling theory and practice; technical but accessible to motivated readers.
- Pew Research Center, “How Public Polling Has Changed in the 21st Century” (2023). Methodologically rich review of how the practice has shifted in the past two decades.
CHAPTER 3
Margin of Error and What It Doesn’t Tell You
“Margin of error: ±3 percentage points.” Almost every published poll carries this notation, often near the bottom of the methodology box. Most readers treat the margin of error as a kind of guarantee — the result is correct, give or take three points. This is wrong, and the misunderstanding is one of the most consequential errors in casual poll reading. The margin of error is a real and useful number, but it measures something narrower than most readers assume, and it ignores the larger sources of error that affect modern polls.
What the margin of error actually measures
The margin of error is a statistical estimate of the uncertainty introduced by random sampling alone. It answers a specific question: if you took many random samples of this size from this population, how much would the results vary across those samples just because of the chance of who got selected?
For a poll with about 1,000 respondents and a 95 percent confidence level, the margin of error is roughly ±3 percentage points. This means: if you repeated the same poll, with the same methodology, many times, in 95 of every 100 repetitions the result would fall within 3 points of the result you got. Five percent of the time, by random chance alone, you would get a result outside that range.
Notice what this number does not include. It does not account for non-response bias (the people who did not answer the phone differing systematically from those who did). It does not account for question-wording effects (the way the question was phrased nudging respondents one direction). It does not account for sampling-frame errors (some kinds of people having lower probability of being included). It does not account for measurement errors (people misunderstanding the question, lying about their views, giving the answer they think the pollster wants). It does not account for any error that comes from a flaw in the design rather than from random variation.
Total survey error: the larger picture
Survey statisticians use the broader concept of “total survey error” to describe all the ways a poll can be wrong. This includes:
- Sampling error. The variation captured by the published margin of error. Pure random sampling variation.
- Coverage error. The error from a sampling frame that does not include all of the target population (the Literary Digest’s 1936 frame missing poorer Americans).
- Non-response error. The error introduced when those who refuse to participate differ systematically from those who do, in ways relevant to the poll’s questions.
- Measurement error. The error from question wording, questionnaire design, interviewer effects, and the gap between what people say and what they actually think or do.
- Processing error. Errors in coding, weighting, and analysis after the data is collected.
Of these, sampling error is typically the smallest in modern polling. The other sources, especially non-response error, are typically larger. A poll with a published margin of error of ±3 might have a true total error of ±5 or more. The published margin gives a sense of the precision of the sampling, not of the overall accuracy of the estimate.
Margin of error in close races
The implications for election polling are substantial. Imagine a poll showing Candidate A at 49 percent and Candidate B at 47 percent, margin of error ±3. Headlines may report “Candidate A leads by 2 points,” but a careful reading is: “within the margin of error, the candidates are statistically tied, with Candidate A possibly anywhere from 46 to 52 percent and Candidate B possibly anywhere from 44 to 50 percent.” In a close race, multiple polls each within the margin of error of “tied” do not collectively prove anything about who is ahead; they prove that the race is close. This is true even before considering total survey error, which makes the picture noisier still.
This is one reason aggregating polls is so important in close races. A single poll showing a 2-point lead is not strong evidence of a 2-point lead; ten polls all showing 1- to 3-point leads in the same direction is somewhat better evidence; ten polls averaging a 2-point lead with confidence intervals computed from the spread of the polls (rather than from individual polls’ sampling error) is better still.
The margin of error of differences
A subtler point: the margin of error of the difference between two candidates’ numbers is larger than the margin of error of either candidate’s number alone. If each candidate’s share has a margin of error of ±3, the margin of error of their difference is roughly ±6. So in the example above, with a published margin of error of ±3 and a 2-point lead, the margin of error of the lead is actually ±6, meaning the lead could be anywhere from −4 (a 4-point trailing position) to +8 (an 8-point lead). The race is not just “close”; it is in fact statistically indistinguishable from a tie or even from a Candidate B lead.
Most reporting does not make this adjustment. Headlines treat the difference between two numbers as if it had the same precision as either number alone, which it does not. A reader who keeps the larger margin of error of differences in mind will be more skeptical of stories about narrow leads than the headline writers.
The 95 percent confidence level: not what most readers think
Most published polls report a 95 percent confidence interval. This is a technical term with a specific meaning that does not match the most natural reading. It does not mean “there is a 95 percent chance that the true value is in this range.” It means: “if we ran this poll many times, 95 percent of the resulting confidence intervals would contain the true value.” The distinction is real and matters for careful interpretation, though for most practical purposes the casual reading is close enough.
More important is the implication: 5 percent of the time, by definition, the published confidence interval does not contain the true value. Across the many polls a typical political junkie reads in any election cycle, some will be wrong by more than the margin of error — not because the pollsters were incompetent but because the math says one in twenty will be. When a single poll shows a striking result that other polls do not show, this is usually noise, not signal.
The bottom line
The margin of error is a statistical estimate of how much sampling variation alone could shift the result — not a guarantee that the published number is within that range of the truth. Real polls have additional sources of error (coverage, non-response, question wording, processing) that are typically larger than sampling error and not included in the margin. The margin of error of a difference between two candidates is larger than the margin of either alone. In close races, single polls within their margins of error are not evidence of a lead; they are evidence the race is close. Treat the margin as a floor on uncertainty, not a ceiling.
What to read or watch next
- Charles F. Manski, Public Policy in an Uncertain World (Harvard University Press, 2013). Manski has written extensively on the gap between published margin of error and total survey error; this book is the accessible synthesis.
- Pew Research Center, “What Surveys Mean” guides at pewresearch.org/methods. Excellent free explainers on margin of error and confidence intervals.
- Andrew Gelman, Statistical Modeling, Causal Inference, and Social Science (statmodeling.stat.columbia.edu). Gelman’s blog is the most consistent source of careful, accessible commentary on polling and statistics in public life.
PART TWO
The Questions and the People
Where the bias hides: in question wording, in who responds, and in the adjustments pollsters make
CHAPTER 4
Question Wording: Where the Bias Hides
In 1941, the pollster Hadley Cantril asked one half of his sample, “Do you think the United States should permit public speeches against democracy?” Sixty-two percent said no. He asked the other half, “Do you think the United States should forbid public speeches against democracy?” Forty-six percent said yes. The same question — should anti-democratic speeches be allowed? — produced opposite-leaning answers depending on whether the survey asked about “permitting” or “forbidding” the same activity. The case became a textbook illustration of how question wording shapes responses, and the underlying lesson has been confirmed thousands of times since: small changes in how a question is asked can produce large changes in how people answer.
The kinds of wording effects
Question wording effects come in several recognizable patterns. A literate consumer of polls learns to spot them:
- Loaded language. Words with strong connotations push respondents toward particular answers. “Do you support the death tax?” produces different responses than “Do you support the estate tax?” — the words refer to the same thing but carry different valences. “Pro-life” and “anti-abortion” describe the same position but with opposite framing. Polls written by advocates often contain such loaded phrases, sometimes obviously, sometimes subtly.
- False dichotomies. Questions that force respondents into an artificial binary when their actual views are more complex. “Do you support cutting Medicare or maintaining it?” may not capture views like “support some changes but not others” or “want to expand certain benefits but reduce others.” Forced binaries simplify analysis but can misrepresent actual public opinion.
- Leading framings. Questions that supply the reasoning for a particular answer. “Given that X causes Y, do you support measures to address X?” guides the respondent toward the answer the question already endorses. Compare to “Do you support measures to address X?”, which leaves the respondent to supply their own reasoning.
- Order effects. The order in which questions are asked can affect responses. Asking about specific concerns (“How worried are you about crime?” “How worried are you about the economy?”) before asking general satisfaction (“Is the country going in the right direction?”) typically produces different general-satisfaction results than asking the general question first. Sophisticated polling experiments often randomize question order to detect and report these effects; many polls do not.
- Acquiescence bias. Some respondents have a tendency to agree with statements regardless of content. “President X is doing a good job” and “President X is doing a poor job” both attract some agreement from the same people. Better-designed questions present both sides of the proposition (“How would you rate the president’s job performance: good, poor, or somewhere in between?”) to dampen this effect.
- Social desirability bias. Respondents are more likely to give answers they think are socially acceptable. They overstate voting, charitable giving, exercise, and religious attendance. They understate drug use, prejudice, and infidelity. The 1989 Virginia governor’s race showed dramatic effects: pre-election polls showed Doug Wilder, the Black Democratic candidate, leading by 9 points; he won by less than 1 point. The gap was attributed in part to white respondents being unwilling to tell pollsters they would not vote for a Black candidate. (The phenomenon was named “the Bradley effect” after a similar gap in the 1982 California race.)
- Don’t-know suppression. Polls that force respondents to choose among given options (rather than offering “don’t know” or “no opinion”) inflate the appearance of strong opinion. Many Americans have no settled view on most policy questions; pushing them to manufacture one in the moment of the survey produces noisy responses that may not reflect any underlying opinion.
Examples from real polling
The polling literature is full of examples where question wording produces dramatically different results. A few well-documented cases:
On health care reform circa 2010: polls asking about “the Affordable Care Act” produced different results than polls asking about “Obamacare,” and both produced different results than polls asking about specific provisions of the law (banning denial of coverage for pre-existing conditions, allowing children to stay on parents’ plans until 26, the individual mandate). Each was a legitimate way to ask about the law; each produced a meaningfully different picture of public opinion.
On immigration: polls asking about “undocumented immigrants” produced more sympathetic responses than polls asking about “illegal aliens” or “illegal immigrants.” The terms refer to the same population; the words shape the response. Pollsters arguing for one or the other terminology often disclose their preference for political reasons; readers should consider the wording carefully when comparing polls.
On abortion: polling on this topic shows particularly dramatic wording effects. Polls asking simply about “support for Roe v. Wade” show different results than polls asking about specific gestational limits, which show different results than polls asking about specific scenarios. Honest polling on abortion typically asks several questions to capture the underlying complexity rather than seeking a single binary answer.
On taxation: “do you support raising taxes on the wealthy” produces different results than “do you support raising taxes on incomes above $250,000” or “do you support raising taxes on people who earn 400 percent more than the median household.” All three describe overlapping populations; the framing matters.
Reading polls with question wording in mind
Some practical principles:
- Find the actual question. Reputable pollsters publish their full question wording in methodological notes. Read the actual question, not the headline’s paraphrase. Reporters often paraphrase questions in ways that drift from the precise wording, sometimes consequentially.
- Compare wordings across polls. When multiple pollsters ask similar questions in different ways, comparing the wordings (and the resulting answers) is more informative than relying on any single poll. Different framings tap different aspects of the same underlying opinion.
- Be skeptical of advocacy polls. Polls commissioned by advocacy organizations, campaigns, or interest groups often use wording that pushes respondents toward the desired answer. This is not necessarily dishonest — the organization may genuinely believe its framing is the correct one — but it does mean such polls should be read alongside polls using other framings.
- Watch for question batteries. A series of related questions may be designed so that earlier questions condition responses to later ones. “How concerned are you about [bad thing]?” followed by “Do you support [policy that addresses bad thing]?” will produce more support for the policy than the policy question alone.
- Look for don’t-know percentages. Polls that report “no opinion” or “don’t know” categories give a more honest picture than polls that force respondents into provided options. A 30 percent “no opinion” rate on a complex policy question is informative; a 0 percent rate suggests the pollster suppressed it.
The high road and the low road
Reputable pollsters take question wording seriously. They test wordings in pilots, run experiments comparing alternative phrasings, publish all of their wording rather than selected highlights, and acknowledge when wording effects are likely. Pew Research Center, the Marquette Law School Poll, the New York Times/Siena College poll, the Wall Street Journal poll, and others operate this way. The American Association for Public Opinion Research’s Transparency Initiative (the AAPOR TI) lists pollsters who commit to disclosure of methodological details, including question wording. Looking for membership in the TI or comparable transparency commitments is a useful signal of polling quality.
Less reputable pollsters — including some campaign pollsters whose work is released selectively, advocacy pollsters with explicit agendas, and “push polls” designed to shape rather than measure opinion — take wording less seriously. “Push polls” in particular are not really polls but a campaign tactic disguised as polling: a caller asks the respondent loaded questions designed to plant negative information about a candidate. The information collected is incidental to the persuasive purpose. Push polls are widely condemned by professional polling organizations but persist.
The bottom line
Question wording substantially shapes the answers polls produce; small changes in wording can produce large changes in results. The kinds of wording effects — loaded language, false dichotomies, leading framings, order effects, social desirability bias, don’t-know suppression — are well-documented and recognizable. Reading polls carefully means finding and comparing the actual questions asked, not relying on headline paraphrases. Reputable pollsters take wording seriously and publish their full methodologies; advocacy pollsters often do not. The same underlying opinion can produce dramatically different poll numbers depending on how it is asked about.
What to read or watch next
- Howard Schuman and Stanley Presser, Questions and Answers in Attitude Surveys (Sage, 1996). The classic study of question wording effects, with extensive experimental evidence.
- Norman M. Bradburn, Seymour Sudman, and Brian Wansink, Asking Questions: The Definitive Guide to Questionnaire Design (Jossey-Bass, 2004). The standard practitioner reference for designing survey questions; readable for non-specialists.
- Pew Research Center, “Methods 101: Question Wording” (pewresearch.org). Concise free overview with examples.
CHAPTER 5
Who Responds, and Who Doesn’t
In 1997, when Pew Research Center began tracking response rates to its telephone polls, about 36 percent of those contacted completed an interview. By 2016, the rate had fallen to about 9 percent. Today, response rates for major telephone polls often sit at 1 to 3 percent. The New York Times/Siena College poll — widely respected and considered among the more rigorous in American polling — reaches a response rate of roughly 1 percent in many cycles. To collect the responses for a single survey, the pollsters place hundreds of thousands of calls. The vast majority of Americans never answer.
This collapse in response rates is the central crisis of contemporary public opinion research. The mathematics of polling assumes that respondents are roughly representative of those contacted; when 99 percent of those contacted refuse to participate, the assumption becomes implausible. The 1 percent who do respond may differ from the 99 percent who do not in ways that matter for the questions being asked. Whether and to what extent this matters is the most consequential methodological question facing the polling industry.
Why response rates fell
Several factors converged. The decline of landline telephones reduced the effective coverage of random-digit dialing; many households now have only cell phones, which are harder and more expensive to reach. Robocalls and telemarketing trained Americans to ignore unknown numbers. Caller ID let people screen calls more effectively. The rise of two-earner households reduced the number of adults home to answer calls during typical polling hours. A general decline in social trust may have made Americans more reluctant to share their views with strangers. Polling organizations responded with longer call windows, more callbacks, and online supplements; the underlying problem persists.
Online polling, which now dominates the industry, has its own response problems. Opt-in panels self-select; only certain kinds of people sign up to take surveys for small incentives. Probability-based online panels (where the panel itself was recruited through a random sample) preserve more statistical validity but typically have low cumulative response rates from the original recruitment to any particular survey. Pew Research Center’s American Trends Panel, one of the highest-quality online panels, often reports cumulative response rates of around 3 percent.
Does it matter who responds?
This is the empirical question. If non-respondents differ from respondents only on characteristics unrelated to the questions being asked, low response rates do not bias the results; they just make samples expensive and slow to gather. If non-respondents differ on characteristics that correlate with the answers, the resulting bias can be substantial.
The empirical evidence suggests that for many questions, non-response bias is relatively modest. Pew Research Center has run experiments comparing high-effort surveys (with extensive callbacks and incentives) to standard surveys; for many topics, the results are similar despite the response rate differences. This suggests that the people who eventually respond after multiple callbacks are not radically different from those who respond on the first try.
But for politically charged questions, especially in elections, the picture is murkier. After the 2016 and 2020 presidential elections, both of which produced larger-than-typical polling errors, AAPOR task forces examined the data and concluded that one likely contributor was differential non-response — specifically, that Trump supporters may have been somewhat less likely to participate in polls than Trump opponents. The hypothesis is plausible: groups distrustful of mainstream institutions, including the media that commissions polls, may be less likely to talk to pollsters; if those groups also lean systematically in one political direction, the resulting samples will be biased.
The 2024 case
Pollsters made substantial methodological changes after 2020 to address the apparent under-representation of Republican voters. The 2024 election was, by most measures, a substantially better year for polling. The American Association for Public Opinion Research’s post-election review found that 2024 polling was more accurate than 2020 or 2016, though it noted that polls still slightly underestimated Republican support relative to Democratic support, suggesting that the underlying differential non-response problem was reduced but not eliminated.
This is a useful illustration of the methodological cycle: errors are detected, methodologies are adjusted, the next cycle is better but not perfect, errors of different kinds are detected, and the work continues. Pollsters in 2026 are still adjusting; the 2028 cycle will reveal whether the current adjustments are durable or whether new problems have emerged.
Beyond political polls: the same problem
The non-response problem affects polls on every topic, not just elections. Surveys of consumer confidence, of religious affiliation, of social attitudes, of health behaviors all face the same falling response rates. The result is that almost any survey-based claim about Americans should be read with the awareness that the respondents may differ from non-respondents in ways the survey cannot fully measure or correct for.
This is one reason careful researchers increasingly supplement surveys with other data sources — administrative records, behavioral data, social media patterns, validated voter files. None of these alternatives is a complete substitute for a well-designed survey, but the combination can help triangulate what surveys alone might miss. Citizens reading polls do not need to do this triangulation themselves, but they should be aware that polling alone is becoming a less complete window on public opinion than it was a generation ago.
The bottom line
Response rates to American polls have collapsed from about 36 percent in the late 1990s to single digits today, often 1–3 percent for major telephone polls. This is the central crisis of contemporary polling. Whether and how much it biases results depends on whether non-respondents differ systematically from respondents in ways relevant to the questions; the evidence is mixed and topic-dependent. Election polling errors in 2016 and 2020 were attributed in part to differential non-response among Republican voters; methodological adjustments improved 2024 accuracy without fully solving the problem. The work continues, and citizens should treat polling-based claims with awareness of this underlying methodological strain.
What to read or watch next
- American Association for Public Opinion Research, “2024 Pre-Election Polling: An Evaluation” (2025). The professional association’s post-election review; the standard reference for understanding 2024 polling performance.
- Scott Keeter et al., “What Low Response Rates Mean for Telephone Surveys,” Pew Research Center (2017). Empirical evidence on how much (and when) low response rates bias results.
- Charles Manski, “Communicating Uncertainty in Official Economic Statistics” (Journal of Economic Literature, 2015). The argument that published margins of error understate true uncertainty by ignoring non-response and other factors.
CHAPTER 6
Weighting and Adjustment
After a poll collects its responses, the work is not done. Almost every published survey applies some form of statistical adjustment — weighting — to make the sample better match the demographic composition of the target population. The basic logic is simple: if your sample has too few young people, give the young people who did respond more weight in the analysis, so that the sample’s effective composition matches the population’s. Done well, weighting corrects for known imbalances in who responded. Done poorly, it introduces new errors or amplifies existing ones. And weighting can fix only what the pollster knows to weight on; it cannot correct for biases on unobserved variables.
How weighting works
A typical national poll, after collecting raw responses, compares its sample to known population benchmarks (typically Census Bureau data) on a set of demographic variables: age, gender, race and ethnicity, education, region, and others. If the sample’s composition differs from the benchmark on any of these variables, the pollster applies weights so that each respondent counts more or less in the final analysis according to how their demographic group is represented.
Example: a poll collects 1,000 responses, but only 8 percent of respondents are aged 18–29, while the Census Bureau says this group is 21 percent of American adults. The young respondents are too few in the raw sample. Weighting addresses this by counting each young respondent more in the analysis — effectively, treating each as 2.5 times as influential as their numbers would otherwise suggest. After weighting, the sample’s effective composition (21 percent young) matches the population’s, and the analysis proceeds.
The mathematics gets more complicated when multiple variables are weighted simultaneously (the standard practice). Modern weighting typically uses techniques like raking or post-stratification that adjust for several variables at once, sometimes including subtle interactions. Bayesian methods (multilevel regression with post-stratification, or MRP) are increasingly used for state-level estimates and other applications.
Weighting on what?
The choice of which variables to weight on is consequential. Standard variables include age, gender, race and ethnicity, education, and region. After 2016, when polls under-represented voters without college degrees in ways that contributed to polling errors, education was added to standard weighting schemes by many pollsters that had not previously weighted on it. The 2020 polling errors prompted further methodological refinement, including discussion of weighting on past vote (asking respondents who they voted for in the previous election and adjusting based on the known actual results) — a technique with both proponents and substantial critics.
The deeper challenge is that you can only weight on variables you measure and have benchmarks for. If your sample under-represents some kinds of voters in ways that do not map onto demographic categories, weighting cannot fix the problem. If political distrust of polling correlates with political leaning in ways the weights do not capture, the resulting sample will be biased even after weighting.
The trade-offs of aggressive weighting
Heavy weighting introduces its own problems. When a small subgroup is up-weighted by a large factor, the influence of any individual respondent in that subgroup grows substantially. If those respondents happen to be unusual within their demographic group, their amplified influence skews the results. Pollsters typically cap the maximum weight applied to any respondent to limit this problem, but the cap itself is a methodological choice.
More fundamentally, heavy reliance on weighting reflects a sample that has substantial imbalances to begin with. A sample that needs minimal weighting because it was already roughly representative is more reliable than a sample that requires aggressive weighting to look representative. The cost of declining response rates is that more pollsters now collect samples that require heavier weighting, with the attendant trade-offs.
Likely-voter screens: a special form of weighting
Election polls face a specific weighting challenge: respondents include both those who will actually vote and those who will not. Polls of registered voters typically over-represent groups less likely to actually turn out (younger voters, lower-propensity voters). Polls aim to weight or screen for likely voters using some combination of self-reported intent (“How certain are you that you will vote?”), past voting behavior (“Did you vote in the last election?”), and demographic models of turnout.
Different pollsters use different likely-voter models, which produces some of the differences between polls of the same election. A poll of registered voters and a poll of likely voters of the same race typically show different results; the gap is the likely-voter model. Aggressive screening typically tightens races and excludes lower-propensity voters; loose screening typically widens the polling sample but may include voters who do not show up. The 2024 cycle saw some pollsters expand their likely-voter universes to include lower-propensity voters who turned out to be important to the result.
How to read polls with weighting in mind
A few practical principles:
- Look for transparency about weighting. Reputable pollsters publish their weighting variables and procedures. Members of the AAPOR Transparency Initiative commit to such disclosure. Less transparent pollsters often have more aggressive weighting, less robust samples, or both.
- Note when weighting changes between polls. A pollster who has changed their methodology between cycles may produce results that look like trends but actually reflect methodological shifts. Read the methodological notes.
- Compare polls with different methodologies. If multiple pollsters using different weighting schemes converge on a similar result, the result is more robust than a single pollster’s number. If they diverge, the divergence is itself informative about how much the methodological choices matter for the conclusion.
- Be especially careful with subgroups. Subgroup results in weighted samples can be unreliable in ways that the overall result is not. A small subgroup heavily weighted up to match the population can be especially noisy.
The bottom line
Almost every published poll applies weighting to make the sample better match the population on observable characteristics. Weighting can correct for known imbalances but cannot fix what it does not measure. The choice of weighting variables is consequential; the addition of education weighting after 2016 and discussion of past-vote weighting after 2020 illustrate how the field evolves. Heavy weighting is a sign of a sample that needed substantial correction, not a guarantee that the correction succeeded. Likely-voter screens are a specific form of weighting for election polls, and different screens produce meaningfully different results from the same underlying data.
What to read or watch next
- Andrew Gelman et al., Bayesian Data Analysis (CRC Press, 2013). The standard treatment of post-stratification and modern weighting.
- Pew Research Center, “Methods 101: How Pew Research Center weights its surveys” (pewresearch.org/methods). Accessible explanation of one major pollster’s actual practice.
- Wei Wang et al., “Forecasting elections with non-representative polls,” International Journal of Forecasting 31 (2015): 980–991. The case study showing that aggressive post-stratification can produce reasonable estimates even from non-probability samples — with caveats.
PART THREE
Election Polls
The most-watched and most-contested form of polling: what likely-voter models do, how trends should be read, and what aggregators can and cannot do
CHAPTER 7
Likely-Voter Models
In every American election, somewhere between 40 and 70 percent of the eligible electorate actually votes, depending on whether the election is presidential, midterm, primary, or local. The rest do not. A poll of all registered voters captures the views of both groups; a poll of “likely voters” attempts to capture only the views of those who will actually turn out. Since the views of the two groups can differ substantially, the choice of who counts as a “likely voter” is one of the most consequential decisions in election polling. Different pollsters use different likely-voter models, and the models account for many of the differences between polls of the same race.
Why likely voters matter
In a typical American election, certain groups are more likely to vote than others. Older voters turn out more than younger voters; voters with college degrees more than voters without; voters with high incomes more than voters with lower incomes; voters who voted in the previous election more than those who did not. These patterns are robust across decades, even as their political effects have shifted. In the mid-twentieth century, the higher-turnout groups (older, college-educated, higher-income) leaned Republican relative to lower-turnout groups; in the contemporary period, the relationship is more complex, with college-educated voters now leaning Democratic in many places while voters without college degrees lean increasingly Republican.
The result is that polls of registered voters and polls of likely voters of the same race can show different results, often by 1–3 points or more. The gap reflects the political differences between groups with different turnout propensities. In close elections, this gap is large enough to determine which candidate appears to be ahead in the poll — even though both polls reflect honest measurement of overlapping but different populations.
How likely-voter models work
Pollsters use various combinations of methods to estimate which respondents will turn out:
- Self-reported intent. “How certain are you that you will vote in the upcoming election?” Respondents who say “definitely” are counted as likely voters; those who say “probably” may be counted with reduced weight; those who say “probably not” or “definitely not” are excluded. The problem: voters routinely overstate their intent. People reliably tell pollsters they will vote at much higher rates than they actually do, partly because of social desirability bias, partly because intent in the abstract is different from intent on the busy day of the election.
- Past voting behavior. “Did you vote in the last election?” Respondents who say yes are weighted more heavily as likely voters. The problem: people overstate this too — typically by 10–20 percentage points compared to actual turnout. Pollsters with access to validated voter files can sometimes check self-reported voting against the public record, which both improves accuracy and reveals which respondents are being honest about their voting history.
- Demographic and behavioral models. Some pollsters combine demographic variables (age, race, education, region) and behavioral variables (interest in politics, attention to the campaign) into models predicting turnout probability. Each respondent is then weighted according to their predicted probability of voting.
- Voter file matching. The most rigorous approach is to draw the sample from a list of registered voters and match each respondent to their actual voting history (publicly available in most states). This allows the pollster to apply turnout probabilities based on actual behavior rather than self-report. The 2024 cycle saw increased use of this approach, which is more expensive but typically more accurate than self-report-based models.
Tight versus loose screens
Pollsters can apply tight or loose likely-voter screens. A tight screen excludes respondents who do not clearly indicate they will vote; a loose screen includes more marginal voters. Tight screens typically produce smaller, more homogeneous samples that lean toward the views of the most reliably engaged voters. Loose screens produce larger, more heterogeneous samples that may include voters who do not show up but also include voters who eventually do.
The choice between tight and loose screens involves a trade-off. Tight screens reduce the chance of polling people who do not actually vote, but at the cost of potentially excluding lower-propensity voters who will turn out. Loose screens capture lower-propensity voters but at the cost of including some who will not vote. The right balance depends on the specific election and the level of campaign-driven turnout among lower-propensity voters; this is itself uncertain in advance.
Different pollsters strike the balance differently, which is one reason polls of the same race can diverge. After the 2024 cycle, in which lower-propensity voters appear to have turned out at higher rates than predicted, some pollsters re-evaluated their screens; the methodological discussion continues.
The 2024 turnout question
The 2024 election produced an interesting test case. Polls in the final weeks of the race generally showed a close contest, with the eventual outcome (Donald Trump’s victory over Kamala Harris in both the Electoral College and the popular vote) consistent with what most polls indicated about the trajectory. Post-election analysis suggested that one of the methodological successes of 2024 was improved coverage of lower-propensity voters, particularly Hispanic and younger male voters, who shifted significantly toward Republican candidates compared to 2020. Pollsters who had widened their likely-voter screens captured some of this shift; those who had not, missed some of it.
This is the cycle of methodological learning. Each election reveals patterns that the previous cycle’s polling did not fully anticipate; pollsters adjust; the next cycle reveals new patterns. The lesson for citizens is not that polls are unreliable, but that they are imperfect instruments adjusting to a changing electorate. Reading polls with awareness of this process — rather than treating each cycle’s polls as authoritative or each cycle’s errors as proof that polls are useless — is the more accurate stance.
What this means for poll consumers
- Look for the likely-voter screen description. Reputable pollsters describe their screen in their methodological notes. The screen description tells you something about who the poll actually represents.
- Compare registered-voter and likely-voter results. Some pollsters report both. The gap between them is informative about how much the screen matters in this race.
- Watch for screen changes mid-cycle. A pollster who tightens their screen as the election approaches may produce results that look like a shift in the race but actually reflect a methodological change.
- Accept that turnout is genuinely uncertain. No likely-voter model is perfect, because turnout depends on factors (the weather on election day, last-minute campaign events, voter motivation that fluctuates) that no model can perfectly anticipate. Polls of likely voters are inherently estimates of an outcome that is not yet determined.
The bottom line
Likely-voter models attempt to estimate which poll respondents will actually turn out to vote, since polls of registered voters typically over-represent groups less likely to vote. The choice of model is consequential; tight and loose screens produce different results from the same underlying data. Self-reported voting intent and self-reported past voting are both unreliable in predictable ways; voter-file matching is more accurate but more expensive. Different pollsters strike different balances, contributing to the variation between polls of the same race. The 2024 cycle highlighted the importance of capturing lower-propensity voters whose turnout shifted the result.
What to read or watch next
- Pew Research Center, “What is a Likely-Voter Model?” (pewresearch.org/methods). Concise explanation of how various models work.
- Robert Erikson and Christopher Wlezien, The Timeline of Presidential Elections (University of Chicago Press, 2012). On how polling at different points in the campaign relates to outcomes.
- American Association for Public Opinion Research, post-election polling reviews (2016, 2020, 2024). The professional reviews of each cycle’s polling errors and the methodological discussion that followed.
CHAPTER 8
Tracking Polls and Movement Over Time
News coverage of campaigns often emphasizes movement: the candidate who is gaining momentum, the candidate who is collapsing, the breakthrough moment that changed the race. Polls are the primary evidence for these narratives. But polling movement is one of the most overinterpreted features of election coverage. Most apparent movement in polls is statistical noise, not actual change in voter sentiment. Real movement does occur, but at smaller magnitudes and on different time scales than campaign coverage suggests. This chapter is about distinguishing real movement from noise.
Why most movement is noise
Recall from Chapter 3 that a poll with a margin of error of ±3 implies that, by random sampling alone, two consecutive polls of an unchanging electorate could differ by 6 points or more. If the candidate was actually at 48 percent the entire time, individual polls might come out at 45, 51, 47, 49, 46, 52 — all consistent with a stable 48 percent, just bouncing within the margin of error. To a careful reader, this looks like noise. To a casual reader following daily poll headlines, the same pattern looks like a series of dramatic swings: the candidate is up, then down, then up again.
This is why looking at single polls in isolation is misleading, especially when comparing one poll to the next. The right reference for movement is the trend across many polls over time, not the difference between any two consecutive polls. The math says that some apparent movement will exist by chance even in a perfectly stable race; only sustained shifts across multiple polls of multiple methodologies are likely to reflect actual movement.
Real movement: when it happens and how to spot it
Real movement in polls does occur, but the patterns are typically clearer than “one poll showed a swing.” Real movements include:
- Convention bumps. Both major-party conventions typically produce a polling bump for the nominated candidate, lasting several weeks before fading. The bumps have shrunk in recent decades but remain real.
- Debate effects. Major debates can shift polls, sometimes by several points, sometimes briefly, sometimes durably. The 2024 cycle saw both: the June debate appeared to durably shift the race; subsequent debates had smaller effects.
- Major news events. Significant scandals, indictments, foreign-policy crises, or running-mate selections can move polls. The size of the effect varies; many “game-changing” events turn out, in retrospect, to have moved polls less than the coverage suggested.
- Late deciders. The final two weeks of a campaign sometimes see real movement among voters who had been undecided. Polls in this period are typically more volatile and the polling errors of close races often come from this period.
- Long-term shifts. Approval ratings, party identification, and broad attitudes shift over months and years in response to underlying conditions. These slow shifts are typically the most consequential; the news cycle tends to under-cover them in favor of dramatic short-term swings that are often noise.
How to read trends
A few practical principles for reading polling movement:
- Look at multiple polls, not single polls. A single poll showing a candidate up 5 points where they had been tied is not strong evidence of movement; multiple polls over several weeks all showing the same direction is stronger evidence.
- Use polling averages or aggregators. As discussed in Chapter 9, sites that average polls and adjust for house effects provide more reliable trend estimates than any single poll.
- Watch for methodological changes. A pollster who changes methodology between two surveys may produce numbers that look like movement but reflect methodology. Reputable pollsters disclose changes; less reputable ones often do not.
- Be skeptical of “momentum” narratives. Coverage often imposes a momentum story on what is actually noise. The candidate said to be “surging” in late October because of a few favorable polls may simply be experiencing the random variation any campaign sees, and the surge may evaporate the next week.
- Look at the underlying fundamentals. Political scientists have long argued that fundamentals — the state of the economy, the incumbent party’s tenure, presidential approval — explain more of election outcomes than the campaign’s short-term moves. Polls that consistently align with the fundamentals are typically more reliable than polls showing dramatic deviations from them.
The special problem of internal polls
Campaigns conduct their own internal polls, which are sometimes leaked or selectively released. These should be read with particular skepticism. Campaigns release internal polls strategically: to shape media narratives, to encourage donations, to demoralize opponents, to claim momentum. Polls released by campaigns favor their candidates more often than not, and the underlying methodology is rarely fully disclosed. When a campaign “shares an internal poll showing the candidate ahead,” the rational interpretation is that the campaign believes releasing this poll serves its interests, not that the poll is necessarily accurate.
The bottom line
Most apparent movement in polls is statistical noise rather than actual shifts in voter sentiment. Real movement does occur — from conventions, debates, major events, late deciders, and long-term shifts — but is best detected across multiple polls over time, not from comparing one poll to the next. Polling averages and aggregators provide more reliable trend estimates than single polls. Coverage often imposes momentum narratives on what is actually noise; the careful reader resists these narratives. Internal campaign polls are particularly suspect: they are typically released strategically rather than because they are accurate.
What to read or watch next
- Robert Erikson and Christopher Wlezien, The Timeline of Presidential Elections (University of Chicago Press, 2012). The empirical study of how polls relate to outcomes throughout the campaign cycle.
- Andrew Gelman, “Moving the Needle on Forecasting” blog posts at statmodeling.stat.columbia.edu. Gelman’s commentary on what does and does not move polls is consistently illuminating.
- Nate Silver, The Signal and the Noise (Penguin, 2012). The general framework for distinguishing real signal from noise in polling and forecasting.
CHAPTER 9
Polling Aggregators and Forecasts
In 2008, a baseball-statistics writer named Nate Silver started a website called FiveThirtyEight that aggregated polls and produced probabilistic forecasts of the presidential election. Silver was right about the outcome of nearly every state, and the success of his approach launched a new genre of polling-based political analysis. By the 2024 cycle, FiveThirtyEight had been spun out and reformed several times (Silver eventually launched his own Silver Bulletin), competitors including The Economist, The New York Times, and others published their own forecasts, and polling aggregators were a standard feature of the political news ecosystem. This chapter is about what aggregators do well, what they do poorly, and how to read their output.
What aggregation actually does
A polling aggregator combines results from multiple polls to produce an estimate of the state of the race. The simple version is a polling average: the mean of recent polls’ results, sometimes weighted by recency or sample size. RealClearPolitics has produced a relatively simple polling average since 2002. More sophisticated aggregators apply additional adjustments:
- Recency weighting. More recent polls count more than older ones, on the theory that opinion can shift over time.
- Pollster quality weighting. Polls from pollsters with better historical track records and more transparent methodologies count more than polls from less reliable pollsters.
- House effect adjustment. Pollsters that consistently lean in one direction relative to the average can have their results adjusted to remove the systematic bias.
- Trend smoothing. Statistical techniques that distinguish actual trend movement from sample-to-sample noise.
- Fundamentals integration. Some forecasts (most famously Nate Silver’s, and the Economist’s) combine polls with non-polling information — economic indicators, presidential approval, party tenure — to produce probabilistic forecasts of likely outcomes.
Why aggregation typically works better than any single poll
The mathematical case for aggregation is straightforward. Each individual poll has random sampling error and possible systematic biases. Averaging across many polls cancels out much of the random error and reduces (though does not eliminate) the systematic biases. A good aggregator with a dozen recent polls in its average has roughly the precision of a single poll several times larger than any of them.
The empirical case for aggregation is also strong. Polling averages have generally been more accurate predictors of election outcomes than single polls of similar timing. Aggregators correctly identified the leader in most state races in most recent cycles, even when individual polls fluctuated. The 2016 case is sometimes cited as evidence against aggregation, but on careful review, the aggregators correctly identified the leader in most states and gave Trump a meaningful (15–30 percent) chance of winning even with Clinton ahead in their forecasts. Trump’s victory was within the bounds of what the forecasts treated as plausible.
Forecasts versus averages: an important distinction
Polling averages and probabilistic forecasts are different things. A polling average is an estimate of the current state of the race; a forecast is a prediction of the eventual outcome. The forecast incorporates the polling average but adds additional uncertainty (“the race could shift between now and election day”), additional information (“fundamentals favor the incumbent party,” “the historical accuracy of polls is X”), and produces a probabilistic statement about who will win.
Probabilistic forecasts are often misread. A forecast that gives Candidate A a 70 percent chance of winning is not predicting that A will win; it is saying that, given current polling and historical patterns, A wins in 70 percent of the simulated scenarios and B wins in 30 percent. A 30 percent chance is not impossible; events with 30 percent chances happen all the time. When the candidate with a 70 percent forecast loses, the forecast was not necessarily wrong; the 30 percent event came up. Distinguishing “forecast wrong” from “low-probability outcome occurred” is one of the hardest things for both forecasters and their readers to do.
The major aggregators
As of 2026, the major polling aggregators include:
- FiveThirtyEight. The site Silver founded; now owned by ABC News. Has gone through methodological transitions; readers should check whether the current methodology is consistent with the historical brand.
- Silver Bulletin. Nate Silver’s post-FiveThirtyEight project. Continues the tradition of detailed methodology and probabilistic forecasts.
- RealClearPolitics. Long-running, simpler average; useful but less methodologically transparent than the more sophisticated aggregators.
- The New York Times Upshot. Combines poll-of-polls analysis with broader political reporting; high-quality methodology, frequent updates.
- The Economist forecast. Built by political scientists G. Elliott Morris and others (with various team changes over time); typically integrates polls and fundamentals.
- Decision Desk HQ, Race to the WH, others. A growing field of aggregators with various methodologies.
How to read an aggregator
- Look at the methodology page. Reputable aggregators describe how they weight polls, what house-effect corrections they apply, what fundamentals they incorporate. A reader who skips the methodology and reads only the topline number is misreading the output.
- Treat the forecast probability as a probability. A 65 percent chance is not a near-certainty; a 25 percent chance is not impossible. Probabilistic forecasts are honest about uncertainty in ways that point predictions are not, but they require probabilistic reading.
- Note the variation across aggregators. Different aggregators using somewhat different methodologies often produce somewhat different forecasts. The variation is itself informative; if all aggregators converge, the conclusion is more robust than if they diverge sharply.
- Distinguish polling averages from forecasts. A polling average shows the current state of the race; a forecast adds predictions about how the race might evolve. Both are useful but different.
- Be aware of “herding.” As election day approaches, some pollsters appear to adjust their methodologies to bring their results closer to the polling average, on the theory that being a clear outlier is reputationally costly. This phenomenon, known as herding, can make the apparent agreement of polls less informative than it looks. Aggregators can partially correct for herding by giving more weight to outlier polls when their methodology suggests they may be capturing real signal.
The bottom line
Polling aggregators combine multiple polls to produce more reliable estimates than any single poll, with adjustments for recency, pollster quality, house effects, and trends. Probabilistic forecasts go further by incorporating non-polling information and producing predictions of eventual outcomes. Forecasts must be read probabilistically: a 70 percent chance is not a prediction of certain victory, and a 30 percent chance happens 30 percent of the time. Different aggregators using different methodologies produce somewhat different forecasts; comparing them is informative. Aggregators have generally outperformed single polls in predicting outcomes, though the 2016 cycle is often misread on this point.
What to read or watch next
- Nate Silver, The Signal and the Noise: Why So Many Predictions Fail — But Some Don’t (Penguin, 2012). The book that explains the philosophy behind probabilistic forecasting; broader than just elections.
- G. Elliott Morris, Strength in Numbers: How Polls Work and Why We Need Them (W.W. Norton, 2022). A defense of polling and aggregation, with attention to recent failures and the path forward.
- Andrew Gelman, “Some Thoughts on Election Forecasting” (various blog posts at statmodeling.stat.columbia.edu). Gelman’s sustained commentary on what aggregators get right and wrong.
PART FOUR
When Polls Go Wrong
The history of polling failures, a closer look at 2016–2024, and the broader question of how statistics can mislead even when they are technically accurate
CHAPTER 10
The History of Polling Failures
Modern public opinion polling has produced some of the most precise predictions ever made about the views of millions of people. It has also produced some of the most spectacular failures in the history of social measurement. The history of those failures is part of the history of the field; each major failure prompted methodological changes, and each cycle of changes was followed, eventually, by new failures of new kinds. This chapter walks through the most important polling failures of the past century and what each taught the field.
1936: The Literary Digest
The case is the most famous in polling history and was discussed in this guide’s Foreword. The Literary Digest mailed ballots to ten million Americans, drawn from telephone directories and automobile registration lists. More than two million were returned. Based on this enormous sample, the magazine confidently predicted that Republican Alf Landon would defeat Franklin Roosevelt, 57 percent to 43. Roosevelt won 61 percent to 37 — the magazine was off by 19 percentage points, on the wrong side. The Digest folded shortly after.
The lesson was that sample size cannot compensate for sample bias. The Digest’s sampling frame — telephone owners and car owners in the depths of the Depression — systematically excluded the poorer Americans who were Roosevelt’s strongest supporters. Among those who could afford telephones and cars, Landon may indeed have led; among the broader electorate, Roosevelt was decisively ahead. The lesson, learned and relearned in every generation since, is that the representativeness of the sample matters more than its size.
In the same election, George Gallup’s newer organization correctly predicted Roosevelt’s victory using a much smaller but more carefully drawn sample of about 50,000. The contrast launched the modern era of probability-based public opinion polling.
1948: Dewey Defeats Truman
Twelve years later, the field of public opinion polling itself failed. All three major polling organizations — Gallup, Roper, and Crossley — predicted that Republican Thomas Dewey would defeat incumbent Harry Truman by a comfortable margin. The Chicago Tribune was so confident in the polls that it printed a famous early edition with the headline “DEWEY DEFEATS TRUMAN.” Truman won. He famously held up the wrong-headed paper for photographers.
Several factors contributed. The pollsters had stopped polling weeks before the election, missing late-breaking shifts toward Truman. They used quota sampling, in which interviewers selected respondents matching certain demographic targets but otherwise had discretion — producing samples skewed toward easier-to-interview Americans, who tended to lean Republican. They had not anticipated that undecided voters would break sharply toward Truman in the final days. The 1948 fiasco prompted the field to shift from quota sampling to probability sampling and to continue polling closer to election day. The methodological reforms that followed shaped polling for the next half-century.
1980: Reagan’s late surge
Polls in the final week of the 1980 presidential race showed Ronald Reagan and incumbent Jimmy Carter in a close contest, with some polls showing Carter slightly ahead. Reagan won by 9.7 percentage points in the popular vote and a 489–49 Electoral College landslide. The shift in the final days — driven by the lone televised debate, the unresolved Iran hostage crisis, and economic anxiety — was larger than polls had captured. The lesson reinforced 1948’s: late-breaking shifts can be substantial, and polls’ ability to capture them depends on continuous polling through election day.
1989: The Bradley Effect
In the 1989 Virginia gubernatorial race, Doug Wilder — the Democratic candidate, who would become the first Black elected governor since Reconstruction — led in pre-election polls by about 9 percentage points. He won by less than 1 point. The gap, attributed to white respondents being unwilling to tell pollsters they would not vote for a Black candidate, became known as the Bradley Effect after Tom Bradley, the Black mayor of Los Angeles whose 1982 California gubernatorial loss showed a similar pattern. The phenomenon prompted concern that polls might systematically understate the appeal of white candidates running against Black candidates due to social desirability bias — white respondents wanting to appear racially tolerant to pollsters even when they were not.
Subsequent research suggested that the Bradley Effect, if it ever existed in its strong form, has weakened or disappeared in more recent decades. Barack Obama’s 2008 election produced no significant Bradley Effect; pre-election polls were broadly accurate about his support. The case remains instructive both for its specific finding and for the broader principle that social desirability bias can affect polls in unexpected ways.
Other notable cases
- The 2008 New Hampshire primary. Pre-primary polls showed Barack Obama with a sizable lead over Hillary Clinton in the Democratic primary; Clinton won. The cause was disputed; possible factors included late-breaking voter shifts following Clinton’s emotional response to a debate question, sampling difficulties unique to early-state primaries with unusual electorates, and possibly some lingering Bradley-Effect-like dynamics.
- The 2014 midterm elections. Polls broadly under-estimated Republican candidates in U.S. Senate races, by about 4 points on average. The cause was probably some combination of differential turnout (Republican voters more likely to actually vote than polls suggested) and possible response biases. The error pattern foreshadowed the larger 2016 errors.
- Brexit, 2016. UK polls before the June 2016 referendum on European Union membership generally showed “Remain” with a slight edge; “Leave” won 52 to 48. The error was within the typical range for polling but on the politically consequential side. Methodological reviews afterward identified turnout differentials and possible response biases as contributors.
- The Israeli 2015 election. Polls showed the Zionist Union party leading or tied with Netanyahu’s Likud; Likud won decisively. Israeli polling has its own methodological challenges and the case is less directly informative for American polling, but it is part of the broader pattern of polling difficulty in close races.
The shape of the lessons
Across these cases, several patterns recur:
- Sample bias is harder to detect than sample size. The 1936 Digest had two million respondents and was off by 19 points. Post-1948 reforms shifted the field toward probability sampling. The contemporary problem of differential non-response is the modern incarnation of the same underlying issue.
- Late shifts can be larger than polls capture. 1948, 1980, and several other cases involved real shifts in the final days that polls did not fully detect. Continuous polling through election day helps but does not eliminate the problem.
- Social desirability bias can hide significant minorities of voters. The Bradley Effect in 1989 and the various Trump-era cases suggest that respondents may sometimes conceal their actual preferences when those preferences are stigmatized in their social or media environment.
- Methodological reforms work, but solve the previous problem rather than the next one. Each round of polling failures produced reforms that addressed the immediate cause; subsequent failures often involved different mechanisms that the reforms had not anticipated. The work is ongoing.
- Close races strain the method. Most polling errors discussed here occurred in close races. In landslides, polls almost always identify the winner; in races within 5 points, polls often get the leader wrong even when the magnitude of error is within typical bounds. This is a structural feature of polling, not a fixable flaw: a method with ±3 to ±5 error cannot reliably identify the leader of a 1-point race.
The bottom line
Polling has produced both spectacular successes and spectacular failures across its century. The 1936 Literary Digest disaster taught that sample bias matters more than sample size; 1948’s “Dewey Defeats Truman” taught the importance of continuous polling and probability sampling; 1980 reinforced lessons about late-breaking shifts; the 1989 Wilder race illustrated social desirability bias. Each cycle of failures prompted methodological reforms. Polling is a work in progress, with each generation’s failures driving the next generation’s improvements; close races are particularly hard for the method, regardless of methodology.
What to read or watch next
- Lindley Gould Crossley and George Gallup, various publications from the 1930s–1950s. The original literature from the founders of modern American polling; useful for historical context.
- Sarah Igo, The Averaged American: Surveys, Citizens, and the Making of a Mass Public (Harvard University Press, 2007). Cultural and intellectual history of how polling came to define how Americans understood themselves.
- AAPOR post-election reviews of 2008, 2012, 2016, 2020, and 2024. The professional association’s assessments of each cycle’s polling performance; the standard references.
CHAPTER 11
2016, 2020, and 2024: A Closer Look
The three most recent presidential elections form a coherent unit in the history of American polling. Each tested polling methodology in particular ways; each prompted methodological reforms; each was followed by a cycle that incorporated those reforms. The 2024 cycle, which produced more accurate polling than 2016 or 2020, was the partial vindication of the work. But the underlying challenges — differential non-response, the difficulty of representing voters skeptical of mainstream institutions — are not solved. This chapter looks closely at all three.
2016: The shock
In the days before the 2016 election, polling averages showed Hillary Clinton leading Donald Trump nationally by 3 to 4 percentage points. Clinton won the national popular vote by 2.1 percentage points. Considered narrowly, the national polling was within typical error bounds; Clinton did indeed win the popular vote, by less than the polls suggested. The polling failure, more precisely, was at the state level. State-level polls in the upper Midwest — Michigan, Wisconsin, Pennsylvania — consistently showed Clinton with leads that proved illusory; Trump narrowly won all three, and the Electoral College tipped on those margins. The state polls in those critical states were systematically wrong in the same direction.
The post-mortem identified several contributing factors. State polls in the upper Midwest typically did not weight on education, and that turned out to matter: voters without college degrees, who broke heavily for Trump, were under-represented in samples that did not weight on education. Late-deciding voters broke disproportionately for Trump in the final days, and many state polls had stopped polling early. There was likely some differential non-response among Trump supporters, who may have been less likely to participate in surveys; the AAPOR post-election review found suggestive evidence of this but could not measure it definitively. Together, these factors produced polling errors of 4–7 points in several critical states, on the wrong side of the result.
The headline narrative — “the polls were spectacularly wrong” — was somewhat overstated. National polls were within typical bounds. Some state polls were quite accurate; others were substantially off. The lesson was less that polls had failed entirely and more that specific methodological choices in specific places had produced substantial errors that mattered for the result. The methodological reforms that followed — most importantly, broader use of education weighting — were targeted at these specific failures.
2020: The repeat
Despite the 2016 reforms, the 2020 cycle produced even larger polling errors, in a similar direction. National polls showed Joe Biden leading Donald Trump by 7–9 points; Biden won the popular vote by 4.5 points. State-level errors were larger still: in Wisconsin, polls showed Biden ahead by an average of about 7 points, and he won by 0.6 points. In Florida, polls showed Biden roughly tied; Trump won by 3.4. Several other states showed similar patterns. The 2020 polling errors were among the largest in modern history, larger in average magnitude than 2016, and, like 2016, they were systematically biased in favor of the Democratic candidate.
The AAPOR post-election review, published in 2021, was a sober document. It examined many candidate explanations and ruled most out. The reviewers could not definitively identify the cause but pointed to differential non-response — specifically, that Trump supporters appeared somewhat less likely to participate in surveys — as the most plausible contributor. The COVID-19 pandemic context complicated matters: polling during the pandemic produced unusual sampling dynamics, and possibly Democratic-leaning Americans staying home during the pandemic answered phone calls more than usual.
The depth of the 2020 errors prompted broader reflection. Some commentators argued that the polling industry had structural problems that methodological tweaks could not solve; others argued that the underlying problems were fixable but required sustained methodological work; still others argued that polling errors were simply a feature of measuring close races in a polarized era. The methodological work continued; some pollsters introduced past-vote weighting (asking respondents who they voted for in the previous election and adjusting based on the known result), while others rejected this approach as introducing its own biases.
2024: The partial vindication
The 2024 election was, by most measures, a substantially better year for polling. National polls showed Donald Trump and Kamala Harris in a close contest, with Trump leading by 1–3 points in the final polling averages; Trump won the national popular vote by 1.5 points and the Electoral College decisively. The American Association for Public Opinion Research’s post-election review concluded that 2024 polling was more accurate than 2020 or 2016, with the typical state-level error within historical norms.
That said, polls in 2024 still slightly under-estimated Republican support, suggesting that the differential-non-response problem identified after 2020 was reduced but not eliminated. The methodological reforms — better coverage of voters without college degrees, expanded likely-voter universes that captured lower-propensity voters, validated voter file matching by some pollsters, more careful handling of new voters and lapsed voters — appeared to have contributed to the improvement. But the systematic direction of the residual error — polls still leaning slightly toward Democratic candidates, on average — suggested that the underlying issue had been mitigated rather than solved.
One particular finding of post-2024 analysis was the importance of capturing lower-propensity voters. Hispanic voters and younger male voters shifted significantly toward Republican candidates compared to 2020. Pollsters whose models gave appropriate weight to these voters captured this shift; those whose models did not, missed it. The lesson was that the electorate continues to evolve, and polling methodologies need to keep pace with electoral change rather than relying on assumptions baked in from previous cycles.
Lessons across the three cycles
- The polling errors were real but recoverable. Each cycle’s problems prompted methodological responses, and the response cycle made things better, even if not perfect. Polling is a maturing field, not a broken one.
- The errors had a consistent direction. 2016, 2020, and 2024 all showed polling errors in the same direction — understating Republican support — even as the magnitude varied. This consistency suggests an underlying methodological challenge rather than random noise.
- The challenge is fundamentally about who responds. Differential non-response, particularly along lines of political distrust, is the most likely underlying cause. Solutions involve outreach, weighting, and methods to reach reluctant respondents.
- Close races stress the method. All three cycles featured close races; the resulting polling errors were politically consequential precisely because the races were close. In landslides, the same methodological errors would have been less visible.
- The headline narratives were sometimes overstated. The popular accounts of “massive polling failures” were not always accurate; national polls were often closer to the actual results than the headlines suggested. The careful reading involves engaging with the specific magnitudes rather than the cultural narratives.
The bottom line
The 2016, 2020, and 2024 cycles form a coherent story: substantial polling errors in 2016, larger errors in 2020 in the same direction, and meaningfully reduced errors in 2024 after methodological reforms. The errors have consistently understated Republican support, suggesting a real underlying methodological challenge — most likely differential non-response among Trump supporters. The improvements between cycles show that polling responds to its failures; the residual errors show that the work is not finished. Reading polls in 2026 with awareness of this history is part of careful poll consumption.
What to read or watch next
- American Association for Public Opinion Research, post-election polling reviews for 2016, 2020, and 2024. The professional reviews of each cycle’s performance; the closest the field has to authoritative assessments.
- Christopher Achen and Larry Bartels, Democracy for Realists (Princeton University Press, 2016). On how voters actually vote, with implications for what polls can and cannot capture.
- Edward Mansfield and Diana Mutz, various papers on the 2016 election and its polling. Among the most careful empirical work on what happened in 2016.
CHAPTER 12
How to Lie with Statistics
In 1954, the journalist Darrell Huff published a slim book called How to Lie with Statistics. It became one of the most successful statistics books ever written, selling more than a million copies and remaining in print continuously for seventy years. Huff’s premise was that the public was constantly being misled by statistics presented honestly in form but dishonestly in framing, and that ordinary readers needed a vocabulary for spotting the manipulation. The specific examples have aged in places, but the underlying patterns Huff identified remain in active use today, by all sides of every political and commercial dispute. This chapter is a guide to the most common patterns.
Misleading averages
“The average household income in this neighborhood is $200,000.” True. The neighborhood has nineteen households earning $40,000 each and one household earning $3,240,000. The average ($200,000) is mathematically correct and substantively misleading; no household in the neighborhood actually earns near the average. The median ($40,000) describes the typical household far better.
The choice between mean (the arithmetic average), median (the middle value), and mode (the most common value) substantially affects what “average” communicates. For symmetric distributions (height, IQ scores within a population), all three are similar. For skewed distributions (income, wealth, household size, web traffic), they diverge sharply. A statistic on “average CEO compensation” that uses the mean is dragged up by a few extreme cases; the median CEO earns much less. A claim about “average” anything should prompt the question: which average, and is the distribution skewed?
The truncated y-axis
A bar chart shows two columns labeled “Sales 2023” and “Sales 2024.” The 2024 bar is twice as tall as the 2023 bar. The headline reads: “Sales soar 100 percent year over year.” Reading the y-axis: 2023 sales were $99 million; 2024 sales were $100 million. The y-axis starts at $98 million and tops at $101 million; the visual difference is dramatic, but the actual change is one percent. The chart is technically accurate — the bars are scaled correctly to the data — but visually misleading.
Truncated y-axes are extremely common in political and commercial communication. They make small changes look large and large changes look enormous. The defense — that one cannot always show the full range from zero — is sometimes valid, particularly when the relevant variation is small relative to the absolute scale. But truncated y-axes should be flagged for the reader; in practice, they often are not. A reader who notices that the y-axis does not start at zero (or at a meaningful baseline) should mentally adjust the visual impression toward what the actual numbers show.
Cherry-picked time frames
“Since [date], the stock market has gained 50 percent.” The choice of starting date often determines whether the claim is impressive. If the start date is the bottom of a recession, almost any subsequent period shows large gains. If the start date is the top of a bubble, almost any subsequent period shows large losses. The claim is true regardless of start date; the meaning depends entirely on what start date is chosen and why.
Cherry-picked time frames are pervasive in political claims. “Since [the speaker’s preferred president took office],” “Since [the speaker’s opponent’s policy was enacted]” — the choice of frame is the rhetorical move. Honest analysis typically uses meaningful natural breakpoints (the start of a presidential term, a major policy change, a recession) and explains why the frame is being used. Suspect framing is silent about the choice.
Confusing correlation with causation
“Studies show that people who eat more chocolate win more Nobel Prizes.” Probably true at the country level: wealthier countries eat more chocolate and produce more Nobel laureates, both because they have more disposable income for both consumption and research. The chocolate is not the cause of the Nobel Prizes; both are effects of the underlying wealth. The classic problem of confusing correlation (two things move together) with causation (one causes the other).
The pattern is constant in popular reporting of social science. “Children raised by [some configuration] have better outcomes” may indicate that the configuration causes better outcomes, or that families with the means and stability to maintain that configuration also tend to produce better outcomes for other reasons, or both, in ratios the simple correlation cannot determine. Distinguishing causation from correlation typically requires either experimental designs (rare in social science) or sophisticated statistical methods that try to identify the causal effect amid the noise. Headlines rarely make these distinctions.
Suspect precision
“The new program reduced unemployment by 3.247 percent.” That last digit may be meaningless. Most economic statistics, after all the measurement error and revision, have effective precision of perhaps a few tenths of a percentage point. A claim of three-decimal-place precision in a context where the underlying measurement is imprecise is a red flag. The precision is a rhetorical device suggesting careful analysis where careful analysis may be absent.
This is also visible in poll reporting. A poll with a margin of error of ±3 percentage points might be reported as showing Candidate A at “52.4 percent.” The .4 is within the noise; reporting it suggests precision the data does not support. Reputable pollsters typically report to a single decimal place; one-decimal precision is honest. Reporting to two or three decimals is usually a sign of either inexperienced reporting or a deliberate impression of precision beyond what the data warrants.
The base-rate problem
“The new test for the disease catches 99 percent of cases.” Sounds impressive. Now: if the disease affects one in 10,000 people, and 100,000 people take the test, the test will identify 9,999 false positives (1 percent of 99,990 healthy people) for every 10 true positives (99 percent of the 10 sick people). The test is more likely than not to be wrong on a random positive. The 99 percent figure was about sensitivity; what mattered for any individual was the predictive value, which depends on the base rate.
The same pattern applies in many domains. Predictions about rare events (terrorist attacks, financial collapses, individual misconduct) often involve high false-positive rates that the headline statistics do not capture. “The model identifies 95 percent of fraudulent transactions” may be useful or useless depending on what fraction of legitimate transactions it also flags. The base rate is essential context.
The semi-attached figure
Huff’s phrase for a number that sounds related to the claim but is technically about something different. “Customers who use our product report 40 percent fewer headaches.” Compared to what — to before they started using the product? to a control group? to the population average? The 40 percent figure is real; what it is 40 percent of is unstated, and the unstated baseline is doing all the work. “Students at our school score 30 percent higher on standardized tests.” Higher than what — the local average, the state average, the national average, the bottom-quartile average? The number is specific; the comparison is fuzzy, and the audience is invited to imagine the most flattering possible comparison.
The misleading graph
Graphs are particularly powerful tools for misleading readers, because the visual impression is processed faster than the underlying numbers. Common patterns include:
- 3D pie charts. The 3D effect distorts the apparent sizes of the slices, making slices in the front look larger than slices in the back of the same actual size. Two-dimensional pie charts are clearer; bar charts are usually clearer still.
- Inappropriate scaling. Charts where the visual scale does not match the data scale (a logarithmic axis presented as linear, or vice versa). The reader processes the visual rather than the labels; misleading scales mislead.
- Selective comparisons. Showing one country/state/group’s data alongside another’s in a way that suggests a comparison the underlying data does not support. The graph implies a relationship; the data does not.
- Misleading icons. Pictograms that scale by image size rather than count. A bar chart showing one group’s growth as a small dollar sign and another group’s as a much larger dollar sign visually exaggerates differences.
The bottom-line skill
The skill of reading statistics critically is, in the end, the skill of asking a few questions consistently:
- What exactly is being measured? The headline often paraphrases; the underlying number measures something specific that may differ from the paraphrase.
- Compared to what? Numbers in isolation are usually meaningless. “40 percent” matters relative to a baseline; a “20-point swing” matters relative to a normal range of variation.
- How was it counted? Definitions matter. “Unemployment” in the U-3 measure is different from U-6; “murder rate” depends on how missing-persons cases are counted; “voter turnout” depends on whether you mean of registered voters or of voting-age population.
- What is the source, and what are the source’s incentives? Statistics produced by parties with stakes in the conclusion deserve more skepticism than statistics produced by neutral observers.
- What does the picture look like with full context? Often, an honest framing produces a much less dramatic conclusion than the headline. The reader’s job is to seek that fuller picture before drawing conclusions.
The bottom line
Statistics can be technically accurate and substantively misleading. Common patterns include misleading averages, truncated y-axes, cherry-picked time frames, confusion of correlation with causation, suspect precision, base-rate problems, semi-attached comparisons, and various forms of misleading graphics. The skill of statistical literacy is the discipline of asking what is being measured, against what baseline, with what definitions, by what source, with what context. Statistical claims that resist these questions are typically the ones that need them most.
What to read or watch next
- Darrell Huff, How to Lie with Statistics (W.W. Norton, 1954). The classic text that defined the genre; still in print, still useful, still funny.
- Charles Wheelan, Naked Statistics: Stripping the Dread from the Data (W.W. Norton, 2013). Modern accessible introduction with contemporary examples.
- Tim Harford, The Data Detective: Ten Easy Rules to Make Sense of Statistics (Riverhead, 2021). The economist and broadcaster’s practical guide for citizens.
- Edward Tufte, The Visual Display of Quantitative Information (Graphics Press, 2001). The classic reference on graph design and the misleading graph; technical but rewarding.
PART FIVE
Beyond Political Polls
Statistics in the rest of public life: issue polls and approval ratings, economic measures, crime and health data, and what each tells us
CHAPTER 13
Issue Polls and Approval Ratings
Election polls receive the most attention, but they are only a small part of public opinion polling. In any given month, dozens of polls measure American views on policy issues, presidential approval, party favorability, institutional trust, and a thousand specific topics from gun control to space exploration. These non-election polls have their own characteristic strengths and weaknesses. They are typically not validated against a clear external benchmark (the way election polls are validated against the actual vote), so their accuracy is harder to assess. They are sensitive to the question-wording problems discussed in Chapter 4. And they often measure shallow opinions that respondents have not deeply considered, rather than the firm convictions the headline numbers might suggest. This chapter is about reading these polls carefully.
Approval ratings
Presidential approval is the most-tracked single measure in American public opinion polling. Gallup has been asking some version of “Do you approve or disapprove of the way the president is handling his job?” since the 1930s. The continuity makes the data uniquely valuable for historical comparison: we can see how Eisenhower’s approval compared to Kennedy’s, how Reagan’s evolved over his two terms, how presidential approval typically responds to economic conditions and major events.
A few patterns are robust across the historical record. Presidents typically begin their terms with relatively high approval (the “honeymoon”), see approval decline over time, sometimes recover late in the term, and end with approval shaped heavily by the state of the economy and major events. Wars and major crises typically produce short-term approval spikes (“rally-around-the-flag” effects) followed by longer-term drag if the crisis persists. Economic conditions are the largest single factor: presidents in growing economies with falling unemployment typically have higher approval than presidents in stagnant or contracting economies, regardless of party.
Reading approval ratings carefully means understanding what they capture. They measure overall sentiment toward the president, conflated across many specific issues. They are sensitive to news coverage; presidents whose names dominate the news in negative ways tend to see approval slip. They are sensitive to party dynamics: in the contemporary polarized era, presidents typically have approval near 90 percent within their own party and near 10 percent in the opposing party, with the small movement coming from the smaller pool of independents and weakly-attached partisans. The dynamic range of presidential approval has narrowed substantially over the past several decades; 1990s spikes of 70+ percent (or troughs in the 30s) are less common, with most modern presidents bouncing between 40 and 50 percent for most of their terms.
Issue polls: shallow versus deep opinion
Issue polls aim to measure public opinion on specific policy questions: “Do you support or oppose the proposed law?” “Is the country going in the right direction?” “How concerned are you about climate change?” These polls capture something real, but the something is often shallower than the numbers suggest.
Decades of research distinguish “non-attitudes” from genuine attitudes. On many policy questions, large fractions of respondents have not previously thought carefully about the issue and have no settled view. When a pollster asks the question, they construct an answer in the moment, often based on cues in the question wording, surface impressions of the issue, or party identification. These answers can be quite unstable: the same respondent asked the same question two weeks later may give a different answer; the same respondent asked the question with slightly different framing may give a different answer. The percentages reported with such confidence may reflect a moment’s construction rather than a deep underlying view.
This is one reason single-question polls on policy issues are often less informative than they appear. A poll showing 60 percent support for some policy may indicate genuine widespread support, or may indicate that the question was framed in a way that produced 60 percent support without much underlying conviction. Cross-validation across multiple framings, attention to whether respondents can articulate reasons for their position, and tracking of stability over time all help distinguish settled opinion from constructed-in-the-moment opinion. Most poll reporting does not engage with this distinction.
Country-direction questions
“Is the country headed in the right direction or wrong direction?” is one of the most-asked questions in American polling. The answers swing dramatically over time and serve as a kind of barometer of national mood. They are also often misread.
The question is interpretively ambiguous. A respondent who thinks the country is going wrong may mean the economy is bad, the political environment is hostile, the culture is changing in unwelcome ways, the international situation is unstable, or any combination. The question taps a general mood but does not specify what the mood is about. Different events can move the same indicator in opposite directions: the same news cycle that improves economic indicators may worsen political ones, leaving the right-direction number stable while the underlying mix has shifted significantly.
The number is often read as a referendum on the incumbent administration, but the relationship is loose. Right-direction numbers can be low under presidents whose specific approval is fairly stable, suggesting the country-direction question taps something broader than the president’s personal performance. The numbers tend to be lower in the contemporary period than in earlier decades; respondents in many polls report dissatisfaction with the country’s direction even during times of relative prosperity, which complicates simple readings.
Issue salience versus issue position
A subtle but important distinction: how much voters care about an issue (salience) is different from what their position on the issue is. A poll showing that 70 percent of Americans support a policy is different in implication from a poll showing that 70 percent of Americans rank the underlying issue as one of their top three concerns. Both are useful but they answer different questions.
Salience matters politically because voters rarely vote on every issue; they vote on the issues that matter to them most. A policy with 70 percent support but low salience may not influence elections; an issue with split support but high salience may be decisive. Climate change is a recurring example: polls show majority support for action on climate change in the abstract, but climate change has historically ranked low in voter priority lists relative to economic concerns, immigration, or health care. The 70 percent support figure is real; so is the limited political salience.
Reading polls carefully includes attention to both dimensions. “What percent supports X?” and “How important is X to voters?” are different questions; the policy implications depend on both.
The asymmetry of intensity
On many issues, opposition is more intense than support, or vice versa. A policy with 60 percent abstract support but with the opposing 30 percent caring intensely while the supporting 60 percent care mildly may be politically dead; the intense minority votes the issue, the mild majority does not. Gun policy in the United States is the classic example: polls have for decades shown majorities supporting various gun-control measures, while gun-rights supporters have voted the issue more reliably than gun-control supporters. The result is policies more aligned with the intense minority than the mild majority.
This asymmetry is hard to capture in headline poll numbers, which typically report support and opposition without measuring intensity. More sophisticated polls include intensity questions (“How strongly do you support/oppose this?”) but the results are rarely highlighted in news coverage. The reader who keeps the asymmetry-of-intensity question in mind — “who cares about this enough to vote on it?” — will get a more accurate sense of political consequences than the support/oppose headline alone.
The bottom line
Issue polls and approval ratings have their own characteristic strengths and weaknesses. Approval ratings are tracked over decades and provide useful historical comparison but capture overall sentiment rather than specific evaluations. Issue polls often measure shallow opinion constructed in the moment of the survey rather than settled conviction. Country-direction questions are interpretively ambiguous. Issue salience is different from issue position; both matter politically. Intensity asymmetries can mean a majority position has less political consequence than the headline numbers suggest. Reading these polls well requires more than reading the headline number.
What to read or watch next
- Philip E. Converse, “The Nature of Belief Systems in Mass Publics” (1964). The classic article on the gap between articulated political opinion and underlying attitude structure. Foundational.
- John Zaller, The Nature and Origins of Mass Opinion (Cambridge University Press, 1992). The most influential modern theory of how survey responses are constructed; technical but rewarding.
- Gallup Polls historical archives (gallup.com). The longest continuous record of American public opinion; useful for historical comparisons of presidential approval and other measures.
CHAPTER 14
Economic Statistics and Their Discontents
“The economy” is one of the most-discussed topics in American public life and one of the most-measured. Unemployment rates, inflation rates, GDP growth, consumer confidence, stock market indexes, housing starts, manufacturing surveys — the federal government and private organizations produce a torrent of economic statistics every month. Each measures something specific; each is sometimes invoked as a stand-in for the broader question of “how is the economy doing?” But the specific statistics often diverge, sometimes substantially, and the choice of which statistic to emphasize is itself a political move. This chapter is about reading economic statistics carefully.
Unemployment: U-3 versus U-6
The headline unemployment rate (U-3) measures the percentage of the labor force that is actively unemployed and looking for work. “Actively looking” is the operative phrase: people who have given up looking for work are not counted in U-3. Neither are people working part-time who would prefer full-time work, nor people who have only marginal attachment to the labor force. U-3 is the official rate published monthly by the Bureau of Labor Statistics and the rate that appears in most news coverage.
U-6, also published monthly by the BLS, measures a broader concept: U-3 plus marginally attached workers (those who want a job and have looked recently but not in the past four weeks) plus people working part-time for economic reasons. U-6 is typically about double U-3. In a healthy economy U-3 might be 4 percent and U-6 about 7–8 percent; in recessions, U-3 might rise to 8 percent and U-6 to 14 percent or more.
Both are legitimate measures of different things. U-3 captures people actively in the labor market and unable to find work; U-6 captures the broader picture of underutilized labor. Critics of U-3 sometimes argue it undercounts labor market slack; critics of U-6 sometimes argue it conflates very different forms of unemployment. The reader who cares about the labor market should be aware that both exist and that the choice of which to cite reflects what aspect of unemployment is in focus. A claim that “unemployment is at a historic low” probably refers to U-3; the same claim about U-6 may or may not hold.
There are also conceptual debates about who counts as “unemployed” in the first place. The labor force participation rate — the percentage of adults either working or looking for work — has declined over recent decades for various reasons (aging population, retirement, disability, changing patterns of school attendance). A falling unemployment rate that occurs because more people have left the labor force is different from a falling unemployment rate that occurs because more people have found jobs. Both can produce identical U-3 numbers; the underlying picture is different.
GDP and its limits
Gross Domestic Product is the total value of goods and services produced in the country in a given period. It is the standard headline measure of economic output and growth. “The economy grew at 3.2 percent” refers to GDP growth (specifically, real GDP growth, adjusted for inflation, annualized from the quarterly figure). GDP is a useful and important measure but should not be mistaken for a complete picture of economic well-being.
GDP captures market activity. It does not capture: unpaid household labor (which is substantial); volunteer work; environmental costs and ecosystem services; the distribution of the income produced (a country whose entire growth went to the top 0.1 percent has the same GDP as one where it spread broadly); changes in the quality of life that are not bought and sold; or many other things. A country can have rising GDP while most households see stagnant or declining living standards, if the income is concentrated. A country can have falling GDP during a war while increasing the long-run sustainability of its economy by reducing pollution. The relationship between GDP and human flourishing is real but loose.
This is why economists and policy analysts increasingly supplement GDP with other measures: median household income (capturing the typical household rather than the total), poverty rates, life expectancy, indices of well-being, environmental measures. None of these is a complete substitute for GDP, but the combination is more informative than GDP alone.
Inflation: CPI versus PCE versus core versus headline
Inflation is the rate at which prices rise over time. It sounds simple. In practice, there are several different official measures of inflation, each capturing slightly different things, and the differences can be politically and analytically significant.
- Consumer Price Index (CPI). The most-cited inflation measure, published monthly by the Bureau of Labor Statistics. Measures price changes for a fixed basket of goods and services purchased by urban consumers. The headline CPI captures everything; the “core” CPI excludes food and energy, which are volatile and can mask underlying trends.
- Personal Consumption Expenditures (PCE). The Federal Reserve’s preferred inflation measure, published by the Bureau of Economic Analysis. Differs from CPI in coverage and weighting; tends to run slightly lower than CPI over time. The Fed’s 2 percent inflation target is specified in PCE terms.
- Producer Price Index (PPI). Measures price changes from the perspective of sellers; sometimes a leading indicator of consumer price changes.
- “True” inflation rates from various private sources. Some commentators argue official measures understate inflation by using methods that mask price increases; the academic and statistical-agency response is that the official measures are designed to capture changes in the cost of maintaining a constant standard of living, and the methods are publicly documented. The debate has political valence; readers should be aware that critiques of official inflation measures often come from particular political directions.
The differences across measures matter. A 4 percent CPI inflation rate may correspond to a 3.3 percent PCE rate, with substantially different implications for what the Federal Reserve should do, what real wage growth has been, and how to interpret economic conditions. Reading inflation news with awareness of which measure is being cited — and what the alternatives would say — is part of careful economic literacy.
Real versus nominal
Almost any economic statistic involving dollars has both nominal and real (inflation-adjusted) versions. “Median household income rose 5 percent” may refer to nominal change (the dollar amount went up 5 percent) or real change (the dollar amount adjusted for inflation went up 5 percent). The two can differ enormously: in a year with 8 percent inflation, a 5 percent nominal raise is a 3 percent real cut.
Honest economic reporting almost always cites real figures or makes clear when nominal figures are being used. Less honest framing exploits the ambiguity. “Household income has reached an all-time high” in nominal terms is essentially always true (because of inflation); the same claim in real terms is much harder to support. Readers of economic claims should look for the words “real” or “inflation-adjusted” and treat dollar comparisons across decades that lack them with skepticism.
Other economic measures and their limits
- Stock market indexes. The S&P 500, the Dow Jones Industrial Average, and the Nasdaq Composite measure the value of particular sets of stocks. They are widely cited as economic indicators, but they capture the wealth of stock owners (a minority of Americans, with ownership concentrated at the top) rather than the broader economic situation. The stock market and the typical household’s economic experience can move in different directions over substantial periods.
- Consumer confidence indexes. Surveys (the University of Michigan Consumer Sentiment, the Conference Board Consumer Confidence Index, and others) measure how Americans feel about the economy. They typically correlate with later economic performance but are not direct measures of economic conditions; they capture sentiment.
- Initial unemployment claims. The number of people newly filing for unemployment insurance, published weekly. A relatively timely indicator of labor-market changes.
- Housing starts and permits. Captures activity in residential construction; a leading indicator of broader economic activity.
- Manufacturing surveys (PMI, ISM). Survey-based indices of manufacturing activity; useful but capture only a portion of the economy and are subject to all the survey-related caveats this guide has covered.
Reading economic claims
- Identify the specific measure. “Unemployment is at a historic low” is a different claim than “median household income is at a historic low.” Both are about “the economy”; they measure different things.
- Look for real vs. nominal distinctions. Especially over time, real figures are usually the meaningful ones. Nominal claims that gloss over inflation are typically misleading.
- Consider distribution. Headline figures often describe averages; the distribution underneath the average matters for what “the economy” means for typical Americans.
- Watch for cherry-picked time frames. As discussed in Chapter 12, the choice of starting and ending dates can dramatically change what the data shows.
- Note the source. Statistics from the Bureau of Labor Statistics, the Bureau of Economic Analysis, and the Federal Reserve are produced by professional civil servants under methodological standards that have been stable across administrations. Statistics from advocacy organizations or political campaigns are typically less rigorous.
The bottom line
Economic statistics measure specific things, and “the economy” is a composite of many specific measures that often move differently. Unemployment has multiple official versions (U-3 vs U-6) measuring different things. GDP captures market activity but not many things that matter for human flourishing. Inflation has multiple measures (CPI vs PCE) that differ meaningfully. Real versus nominal distinctions are essential for comparisons over time. Stock indexes capture asset values for owners, not the broader economic experience. Reading economic claims well means identifying the specific measure, looking for distribution, and being skeptical of cherry-picked framings.
What to read or watch next
- Bureau of Labor Statistics (bls.gov). The official source for unemployment, inflation, and other labor-related statistics. Methodology is publicly documented.
- Bureau of Economic Analysis (bea.gov). The official source for GDP and personal consumption expenditures. Methodology is publicly documented.
- Federal Reserve Economic Data, FRED (fred.stlouisfed.org). The single best free source for time series of nearly all economic statistics; built by the St. Louis Federal Reserve. Indispensable for serious engagement with economic data.
- Diane Coyle, GDP: A Brief but Affectionate History (Princeton University Press, 2014). On the strengths and limits of GDP as a measure of economic activity.
- Tim Harford, Fifty Things That Made the Modern Economy (Riverhead, 2017). Engaging short essays on economic measurement and concepts.
CHAPTER 15
Crime Statistics, Health Statistics, and Other Counts
Beyond polls and economic statistics, citizens encounter constant claims based on counts: how much crime there is, how many people died from this or that cause, how many were homeless, how many were enrolled in this program. These statistics shape political debate, policy, and individual decisions about where to live and how to assess risk. They are also frequently misunderstood. The headline numbers often mask substantial measurement issues, and the same underlying reality can produce very different statistics depending on how it is counted. This chapter is about reading these counts carefully.
Crime statistics: UCR versus NCVS
There are two major sources of national crime statistics in the United States, and they often disagree. Understanding why is essential for reading crime claims.
- Uniform Crime Reporting (UCR), now the National Incident-Based Reporting System (NIBRS). The Federal Bureau of Investigation’s collection of crimes reported to police, compiled from state and local agencies. Produces the headline statistics on murders, robberies, assaults, and other reported crimes. The transition from UCR’s older summary system to NIBRS’s incident-based system is ongoing and creates comparability challenges across years.
- National Crime Victimization Survey (NCVS). The Bureau of Justice Statistics’ survey of American households, asking respondents whether they were victimized by crime during the previous period. Captures crimes that may or may not have been reported to police.
The two measures often diverge because they measure different things. UCR/NIBRS captures only crimes reported to police; NCVS captures both reported and unreported crimes. Reporting rates vary by crime type: murder is essentially always reported; robbery and assault are reported in roughly half the cases or fewer; rape and sexual assault have especially low reporting rates, with NCVS capturing many more incidents than UCR/NIBRS. The result is that NCVS-based estimates of crime are typically higher than UCR/NIBRS-based estimates for many categories, and the gap has implications for both how much crime there is and what trends look like.
Reporting rates also vary by jurisdiction, demographic, and historical period. Trends in reporting can mimic trends in crime: if reporting rates rise, the apparent crime rate rises even if actual crime is steady. This complicates simple readings of UCR-based crime trends. Most scholars believe the broad strokes of crime trends are real — the long decline from the early 1990s through about 2014 reflected real declines in crime, not just reporting changes — but specific year-to-year movements should be interpreted carefully.
The 2020s crime debate
The early 2020s saw a particularly contested crime statistics debate. Murder rates rose substantially in 2020 — the largest one-year percentage increase in modern U.S. history — with the trend continuing in 2021. By 2022–2024, murder rates began declining again, returning to roughly pre-pandemic levels in many cities by the mid-2020s. Different commentators emphasized different parts of this trajectory. Supporters of one set of policies emphasized the recent declines; supporters of opposing policies emphasized the higher levels relative to a few years earlier. Both could cite accurate UCR-based statistics; the framing of the time series determined the impression.
This is a useful illustration of the broader pattern. Crime statistics over short windows are noisy; the choice of time frame substantially shapes the apparent picture; multiple legitimate statistics exist for the same underlying phenomenon; and the political stakes of crime statistics ensure that all sides will cite the framings most favorable to their positions. The careful reader looks at multiple framings, multiple sources, and longer time horizons before drawing strong conclusions.
Health statistics: causes, attribution, and excess deaths
Health statistics raise their own measurement challenges. “Deaths from X” depends on what is counted as a cause: a person who dies from pneumonia after years of smoking-related lung damage may be counted as dying from pneumonia, from lung disease, from smoking, or all three, depending on the framework. The official cause of death recorded on a death certificate may differ from what an epidemiologist would consider the underlying cause.
This was particularly visible during the COVID-19 pandemic. The official COVID death count depended on definitions: deaths from COVID, deaths with COVID, deaths where COVID was a contributing factor. Different jurisdictions used different definitions; comparison across countries was complicated by reporting differences. Many researchers argued that the most reliable measure was “excess deaths” — the number of deaths in a period beyond what would have been expected based on prior years — which captures both direct COVID deaths and pandemic-related indirect deaths (delayed care, mental health impacts, and so on). The excess deaths framing produced estimates that were typically substantially higher than the official COVID death counts in many countries.
More broadly, health statistics often involve trade-offs between specificity (what specific cause this person died from) and accuracy (what would have happened in a counterfactual world without this risk factor). Both matter; neither alone is fully informative.
The base rate problem in risk reporting
Headlines often report relative risks (“X raises your risk by 50 percent”) without the base rate that gives the relative risk its meaning. “50 percent higher risk of cancer” is alarming if the base rate is 20 percent (raising it to 30 percent) and almost meaningless if the base rate is 0.001 percent (raising it to 0.0015 percent). The relative risk is the same; the implications for an individual’s decision are entirely different.
Honest health journalism includes both relative and absolute risks. Less honest framing emphasizes whichever framing produces the more dramatic impression, which depends on context. Pharmaceutical advertising sometimes emphasizes absolute risks (“only 1 in 10,000 patients had this side effect”) while critics emphasize relative risks (“doubled the risk of this side effect”); the same data supports both framings, with very different impressions.
Counts that depend on definitions
- Homelessness counts. Estimates of how many Americans are homeless depend on definitions (people in shelters, people unsheltered, people doubled up with friends or family due to lack of housing) and on counting methodologies (point-in-time counts on a single night versus annual counts that include all who experienced homelessness in a year). Estimates can vary by factors of 5 or 10 depending on the definition.
- Poverty counts. The official U.S. poverty measure uses a methodology developed in the 1960s that excludes many forms of assistance (SNAP benefits, housing assistance, the Earned Income Tax Credit). The Supplemental Poverty Measure, developed more recently, includes these and produces somewhat different estimates. Both are official; both measure something meaningful; they often disagree about specific subgroups.
- Education statistics. High school graduation rates depend on whether students who take five years to graduate count as graduates or not. College enrollment depends on whether part-time and online students are counted equally. Test score comparisons depend on which test, which year, and what populations are included.
- Immigration statistics. The number of unauthorized immigrants in the United States depends on definitions and estimation methods; estimates from credible sources vary by several million. The number of asylum seekers depends on what stage of the process counts. The number of illegal border crossings depends on whether counts include attempts that were turned back versus successful crossings.
How to read counts and rates
- Identify the specific definition. “Murders” and “homicides” are different (homicides include justifiable killings); “unemployment” differs by U-3 versus U-6; “poverty” differs by official versus supplemental. The headline label may obscure the specific definition.
- Identify the source. Federal statistical agencies (BLS, BEA, FBI, CDC, BJS, Census Bureau) typically follow rigorous documented methodologies. Advocacy organizations’ statistics may or may not. Both can be useful, but the source affects how skeptically to read.
- Look for relative versus absolute risks. Especially for health and crime claims, ask both “what is the change?” and “what is the baseline?”
- Beware comparison across definitions or time periods. Apparent changes in counts may reflect definitional or methodological changes rather than changes in the underlying reality. The transition from UCR to NIBRS, the redefinition of poverty, the inclusion or exclusion of specific groups in employment statistics all can produce apparent trends that are artifacts of methodology.
- Use multiple sources. Where possible, look for multiple credible sources estimating the same quantity. Convergence is reassuring; divergence is informative about how much the methodological choices matter.
The bottom line
Counts and rates often involve substantial measurement issues that headline numbers obscure. Crime statistics from UCR/NIBRS and NCVS measure different things and often diverge; the same is true for health statistics with different death-attribution rules; counts of the homeless, the impoverished, and the educated all depend on definitional choices that can change estimates by factors of two or more. Relative risks are meaningful only with base rates. Comparing across time periods is complicated when methodologies have changed. The careful reader identifies the specific measure, the source, the definitions in use, and the trade-offs involved before drawing conclusions about “how much” of anything there is.
What to read or watch next
- Bureau of Justice Statistics (bjs.ojp.gov). The federal source for crime victimization data and the standard for criminal-justice statistics.
- FBI Crime Data Explorer (cde.ucr.cjis.gov). The interactive interface for FBI crime data; useful for exploring trends and comparisons.
- Centers for Disease Control and Prevention (cdc.gov). The federal source for U.S. health statistics, including mortality, morbidity, and risk factors.
- Hannah Ritchie, Not the End of the World (Little, Brown Spark, 2024). The Our World in Data researcher’s primer on reading health and environmental statistics with appropriate skepticism and context.
- Andrew Gelman and Deborah Nolan, Teaching Statistics: A Bag of Tricks (Oxford, 2017). For readers wanting to engage more deeply with how statistical thinking works in practice.
PART SIX
Putting It Together
A practical toolkit, what polls can and cannot tell you, and the civic stakes of statistical literacy
CHAPTER 16
A Citizen’s Practical Toolkit
This chapter pulls together the practical work of being a careful consumer of polls and statistics into a single set of tools. The previous chapters laid out the underlying concepts; this one offers checklists and habits that turn the concepts into practice. Statistical literacy is not a static body of knowledge but a set of habits of attention. The habits become natural with practice.
The five-question checklist for any poll
When you encounter a poll claim in news, on social media, or anywhere else, run through these five questions before drawing conclusions:
- 1. Who conducted the poll? Reputable pollsters with track records of transparency (members of the AAPOR Transparency Initiative, established academic and media polling operations, longstanding industry players like Pew, Gallup, Marist, Marquette, NYT/Siena, ABC/Washington Post, Wall Street Journal) deserve more initial trust than unknown pollsters, advocacy organizations, or campaign-affiliated pollsters. Look for the name of the polling organization in the article; if not provided, that itself is informative.
- 2. What was the actual question? Reputable polls publish their full question wording. Read it. The headline’s paraphrase often drifts from the actual question in ways that affect interpretation. Loaded language, leading framings, and false dichotomies are easier to spot in the actual question than in the paraphrase.
- 3. Who was sampled, and how? The sampling frame (likely voters, registered voters, all adults, particular subpopulation) affects what the poll can be used for. The methodology (probability sample, opt-in panel, online versus phone) affects how reliable the sample is likely to be. Polls without methodology disclosed are typically less reliable than polls with it.
- 4. What is the margin of error — and remember its limits. The margin gives you the random sampling uncertainty. Real polls have additional sources of error not captured in the margin. The margin of differences between candidates is roughly twice the margin of either alone. In close races, polls within their margins of error are evidence the race is close, not evidence of a clear leader.
- 5. How does this fit with other polls? Single polls are noisy. The relevant question is whether this result is consistent with other polls of the same race or topic, or whether it stands out. Outliers exist, but a single outlier deserves more skepticism than a result that aligns with broader patterns.
The five-question checklist for any statistical claim
Beyond polls, statistical claims of all kinds appear constantly. Run through these:
- 1. What exactly is being measured? Often the headline term (“crime,” “unemployment,” “poverty”) covers multiple distinct measures. Identify which one is being cited.
- 2. Compared to what? Numbers in isolation usually do not mean much. “Rose 5 percent” matters relative to typical variation; “saved 100 lives” matters relative to the total at risk; “Historic high” matters relative to whether the underlying baseline is appropriate.
- 3. What time frame is being used, and why this one? Cherry-picked time frames are a primary tool of misleading framing. A claim that uses a particular start or end date should be tested by trying alternative reasonable dates.
- 4. What is the source, and what are its incentives? Statistics from neutral statistical agencies are typically more reliable than statistics from organizations with stakes in the conclusion. A statistic produced by a party that benefits from the conclusion deserves more scrutiny than one produced by a neutral observer.
- 5. What does the picture look like with full context? Often, the honest framing of the same data produces a much less dramatic conclusion than the headline. Seeking the fuller picture before drawing conclusions is the central discipline.
Habits to cultivate
- Read multiple sources before forming views. Any single source has its angle; cross-reading reveals the angle and the underlying ground. This is true for political news, for poll reporting, for statistical claims of all kinds.
- Pay attention to the methodology. Polls and statistics that omit methodology are typically less reliable than ones that include it. Reading the small print is part of careful consumption.
- Distinguish description from prediction. A poll showing the current state of a race is different from a forecast predicting the outcome. Both can be useful; conflating them produces errors.
- Maintain calibrated uncertainty. Most statistical claims involve more uncertainty than they convey. A range of plausible outcomes is usually more accurate than a point estimate. Treating point estimates as if they were certainties is one of the most consistent sources of statistical mistakes.
- Update on evidence. New polls, new data, new analyses should update your views. Holding onto a position that the evidence has shifted away from is a failure of intellectual honesty.
- Be skeptical of dramatic claims. The headline that promises a stunning finding usually delivers something less stunning when read carefully. Calibrate your initial credence to be lower for dramatic claims and adjust upward only with strong evidence.
- Resist the urge to share before reading. Social media incentives reward fast sharing of striking claims; the careful reader resists the urge to share until the claim has been verified.
- Develop favorite sources. Building a stable of trusted sources — polling organizations, statistical commentators, statistical agencies — reduces the cognitive load of evaluating every claim from scratch. The trust should be calibrated to track records, not partisan alignment.
Sources to bookmark
A practical starter set of sources for following polls and statistics carefully:
- Pew Research Center (pewresearch.org). Among the most rigorous American pollsters; their Methods 101 series and ongoing methodological transparency make them an essential reference.
- AAPOR (aapor.org). The professional association of public opinion researchers; post-election reviews, transparency standards, and resources for journalists and the public.
- Silver Bulletin and similar aggregators. For election cycles, polling aggregators with transparent methodologies provide better estimates than any single poll.
- FRED (fred.stlouisfed.org). The St. Louis Fed’s economic data archive; the single best free source for economic time series.
- BLS, BEA, FBI, CDC, Census Bureau. The federal statistical agencies; primary sources for the underlying data behind most economic, crime, health, and demographic claims.
- Statistical Modeling, Causal Inference, and Social Science (statmodeling.stat.columbia.edu). Andrew Gelman’s blog; the most consistent source of careful, accessible commentary on polls and statistics in public life.
The bottom line
Statistical literacy is a set of habits of attention, not a static body of knowledge. The five-question checklists for polls and statistical claims provide practical structure: who conducted, what was asked, who was sampled, what is the uncertainty, how does this fit with other evidence; what is being measured, compared to what, over what time frame, by what source, with what context. Cultivating habits of multi-source reading, methodology attention, calibrated uncertainty, and resistance to dramatic claims makes the work easier over time. A small stable of bookmarked, trusted sources reduces cognitive load.
What to read or watch next
- Tim Harford, The Data Detective: Ten Easy Rules to Make Sense of Statistics (Riverhead, 2021). The most readable practical guide to evaluating statistical claims as a citizen.
- Carl Bergstrom and Jevin West, Calling Bullshit: The Art of Skepticism in a Data-Driven World (Random House, 2020). Companion to a popular University of Washington course; specific tools for spotting statistical manipulation.
- Charles Wheelan, Naked Statistics: Stripping the Dread from the Data (W.W. Norton, 2013). Accessible introduction to statistical concepts for the curious general reader.
CHAPTER 17
What Polls Can and Cannot Tell You
After fifteen chapters on the mechanics, history, and practical use of polls and statistics, this chapter takes a step back and asks the more philosophical question: what can these instruments actually tell us, and what are their fundamental limits? A careful citizen needs both ends of this answer. Polls do real work; treating them as worthless is its own form of error. They have real limits; treating them as authoritative is the more common error. Honest engagement involves accepting both.
What polls can tell us
- The rough state of public opinion. A reasonably well-conducted poll on a relatively settled question can provide a reasonable estimate of what the population thinks. The estimate has uncertainty, the methodology matters, the question wording matters — but, with care, polls can give us a defensible picture of where opinion sits.
- The rough state of an election race. Aggregating multiple election polls produces estimates of the standing of candidates that have generally outperformed both individual polls and informal impressions. The estimates are imprecise in close races but are informative about whether the race is close or not.
- Trends over time. Sustained shifts in opinion across many polls, especially when consistent in direction across many pollsters, are real signal. The decline in public confidence in major institutions over the past several decades, the long shifts in attitudes toward racial integration or gay marriage, the cyclical pattern of presidential approval — all of these are visible in polling data and meaningful.
- Differences between groups. Polls can reveal real differences in views between demographic groups, between regions, between political subcultures. These differences may or may not be the most politically salient features of the moment, but they are typically real.
- The state of the economy and other measurable conditions. The federal statistical agencies’ measures of unemployment, inflation, GDP, crime, and so on, while subject to all the caveats this guide has discussed, do measure something real about the economic and social state of the country. Reading them carefully provides genuine information about how things are going.
What polls cannot tell us
- What people will think next month. Polls are snapshots. They cannot predict how opinion will evolve, especially in response to events that have not yet happened. Forecasts based on polls add additional information but still capture only what is known now; surprises remain possible.
- How intensely people care. Standard polls capture position but not intensity well. The political consequences of public opinion depend heavily on intensity, and intensity is often distributed asymmetrically (the minority that intensely opposes a policy may matter more politically than the majority that mildly supports it). Polls rarely capture this well.
- Why people hold the views they do. Polls capture what people will say in response to a question; they do not capture the underlying reasoning, the experiences that shaped the view, or the conditions under which the view might change. For these, deeper qualitative methods (interviews, ethnography, focus groups) are necessary, and even these have their limits.
- What the right policy is. Polls are descriptions of what the public thinks; they are not normative arguments about what should be done. Even if 75 percent of Americans support a policy, the policy could be wrong; the public has been wrong about important questions in the past, and respect for democracy does not require treating poll majorities as oracles. Polls inform political decisions; they do not settle questions of justice or wisdom.
- How people will actually behave. Self-reported intent is often a poor guide to behavior. “I plan to vote” is not the same as “I will vote.” “I would consider buying this product” is not the same as “I will buy this product.” Behavioral data (actual votes, actual purchases) is typically more reliable than survey-based behavioral predictions.
- What is true about the world. Polls measure what people believe; they do not measure what is true. A poll showing that 60 percent of Americans believe X is informative about beliefs, not about X. Polling on contested factual questions (“do vaccines cause autism,” “was the 2020 election fair”) reveals the social distribution of beliefs but does not adjudicate the questions themselves. The reader who confuses widespread belief with truth makes an error of category.
The deeper limits
Beyond these specific points, polling has several deeper limits that are worth pausing on:
The limits of any single moment
A poll captures a moment. Public opinion is not a static thing; it shifts in response to events, to coverage, to social discussion, to argument. The poll that shows 60 percent support today may show 50 percent support in three months, not because anyone in particular has changed their mind decisively but because the ground itself has shifted. Treating any single poll — or even any single moment’s aggregate of polls — as a fixed measurement of “what the country thinks” obscures the dynamic, evolving nature of public opinion.
The limits of asking strangers
Polls measure what people will say to a stranger asking questions on the phone or online. This is not the same as what people privately believe, what they tell their friends, or what shapes their actual decisions. Some views are stigmatized in ways that make people reluctant to share them with strangers; some views are constructed in the moment of the survey rather than being held as settled positions; some views are loosely held and easily moved by question wording. The poll number is the result of asking strangers a question; it is not a transparent window into the underlying opinion of the population. The two are related, but the relationship is loose.
The limits of reducing complex views to single numbers
Most political and social questions are not reducible to a single approve/disapprove or yes/no. Real views on most issues are complicated, contextual, and often internally tension-laden. People often think that some aspects of a policy are good and others are bad; that a politician has done some things well and others badly; that an institution serves some functions well and others badly. Polls reduce these complex views to single numbers, which is necessary for measurement but loses substantial information. The reduction is a useful approximation, not a complete picture.
The limits of measurement-driven politics
Beyond the technical limits, there is a broader concern about what intensive polling does to democratic deliberation. When public opinion is constantly measured and reported, politicians may respond to the measurements rather than leading on the merits; voters may form views based on what they perceive other voters to think, rather than on independent reflection; the daily horse-race coverage of campaigns may crowd out substantive engagement with the actual choices at stake. These concerns do not invalidate polling — the alternative of governing with no idea what citizens think has its own problems — but they suggest that polling is part of the political environment, not a neutral observer of it.
A balanced stance
The right stance toward polls and statistics is neither credulity nor dismissal but calibrated engagement. Polls can be informative when read carefully; they can be misleading when read poorly. The work of citizenship in a polled democracy is to develop the habits of careful reading, while also recognizing that polls are one source of information among many and that ultimate political judgments rest on more than what polls report.
This guide has tried to provide the practical tools for that careful reading. The tools are not magic, but they help. The reader who has worked through these chapters — who can spot a loaded question, who can read a margin of error correctly, who can distinguish description from prediction, who knows what the major sources are and how to find them — is better equipped to engage with public claims than the reader who has not. That improvement is small relative to the larger problems of contemporary public life, but it is real, and it accumulates.
The bottom line
Polls can tell us the rough state of public opinion, the rough state of election races, trends over time, differences between groups, and the state of measurable conditions. They cannot reliably tell us what people will think later, how intensely they care, why they hold the views they do, what the right policy is, how people will actually behave, or what is true about the world. The deeper limits include the limits of any single moment, the limits of asking strangers, the limits of reducing complex views to single numbers, and the broader effects of measurement-driven politics. The right stance is calibrated engagement: neither credulity nor dismissal, but careful reading.
What to read or watch next
- Walter Lippmann, Public Opinion (1922). The classic work on the limits of public opinion as a basis for governance. Still essential a century later.
- Susan Herbst, Numbered Voices: How Opinion Polling Has Shaped American Politics (University of Chicago Press, 1993). Historical and theoretical reflection on what polling has done to American democracy.
- Justin Lewis, Constructing Public Opinion: How Political Elites Do What They Like and Why We Seem to Go Along with It (Columbia University Press, 2001). A more critical perspective on the production and use of poll numbers.
CHAPTER 18
The Civic Stakes of Statistical Literacy
This guide closes with the question of why statistical literacy matters — not just to individuals trying to read the news intelligently, but to the larger project of self-governance in a republic. The case is broader than “it will help you spot misleading claims,” true as that is. Statistical literacy is part of the equipment of citizenship in any modern democracy, and its erosion has consequences for the quality of self-governance itself.
Democracy depends on accurate measurement
Representative democracy assumes that the views and conditions of citizens can be, in some rough way, known and represented. Elected officials respond to the conditions and preferences of the public; voters evaluate officials based on outcomes that can be measured. Both depend on functioning measurement of public opinion, economic conditions, social outcomes, and the like. When the measurements break down — when polls cannot accurately capture public opinion, when economic statistics are routinely contested, when health and crime data are systematically misread — the link between citizen and government strains.
This is not a hypothetical concern. The polling failures of 2016 and 2020 contributed to widespread skepticism about whether polls could be trusted at all. Disputes over economic statistics (“real” inflation versus the official rate, “real” unemployment versus U-3) have become routine. Different sources of crime statistics support different political narratives. The result is an information environment in which agreement on the basic facts of public life has become harder to achieve, and in which the work of policy debate is preceded by, and sometimes substituted by, disputes about what the facts even are.
The asymmetry of the problem
A particularly difficult feature of the contemporary statistical environment is that the costs of misreading and the costs of distrust are asymmetric. A citizen who naively trusts a misleading poll is wrong; a citizen who reflexively distrusts all polls is also wrong. But the two errors are not symmetric in their political effects. Naive trust enables manipulation; reflexive distrust enables a different kind of manipulation, in which dismissing inconvenient evidence becomes acceptable because “all statistics are misleading.” Both errors damage democratic discourse, but the second error has been growing in recent years and has particular dangers.
The careful path is calibrated engagement: trust evidence proportionate to its quality, recognize the limits of any specific measurement, but maintain the basic conviction that careful measurement is possible and that the available data, properly read, is a real source of information about the world. This is harder than either extreme. It requires sustained intellectual work. But it is the precondition of evidence-based political deliberation, and the alternatives are worse.
Statistical literacy and political polarization
Polarization makes statistical literacy harder, in a particular way. When statistics are perceived as politically aligned — “those are Democratic numbers,” “those are Republican numbers” — the work of evaluating them is contaminated by political loyalty. Citizens of one party may dismiss statistics that contradict their preferred narrative without engaging with the substance; citizens of the other party may do the same. Both are responding to perceived political alignment rather than to methodological quality. The result is a loss of the basic capacity to engage with shared evidence.
The federal statistical agencies — the Bureau of Labor Statistics, the Bureau of Economic Analysis, the Census Bureau, the FBI’s crime statistics, the CDC — are political only at the margin. They are run by professional civil servants under methodologies that have been stable across administrations of both parties; their data is published according to predetermined schedules; their methodologies are publicly documented. The data is not perfect, but it is substantially less politicized than the rhetoric around it suggests. Citizens who develop the habit of treating these sources as common ground — even when interpreting them in different ways — are doing something important for democratic discourse. Citizens who treat them as partisan instruments contribute to a more dangerous form of polarization.
Statistical literacy and the information environment
Contemporary information environments make statistical literacy both harder and more necessary. Social media spreads striking statistical claims at high speed; the incentive structure rewards what is dramatic and shareable, not what is accurate; misinformation can travel faster than corrections; the volume of statistical claims is far greater than any individual can carefully evaluate.
The response is not to give up but to develop the discipline of selective engagement. A citizen who carefully reads a few statistical claims a week, returning to original sources and primary methodology, is doing more for their understanding than a citizen who reflexively shares dozens of claims daily without verifying any. The work is, paradoxically, lighter and more durable when done carefully on fewer claims than when attempted at the volume the information environment seems to demand.
Statistical literacy as civic responsibility
This guide has framed statistical literacy as a practical skill, but it is also a civic responsibility. In a republic, citizens are not just consumers of information but participants in the public conversation. The quality of the public conversation depends on the quality of the participants. Citizens who carefully read polls and statistics, who refuse to spread claims they have not verified, who distinguish honest measurement from partisan manipulation, who hold their own preferred sources to the same standards they apply to opposing ones — these citizens raise the level of public discourse by their participation. Citizens who do the opposite lower it.
This is not a heroic burden; it is the ordinary work of citizenship in a country that depends on the judgment of its citizens. Most of the work happens quietly, in private reading, in conversations with family and friends, in decisions about what to share and what to verify before sharing. The aggregate effect of millions of citizens doing this work or not doing it is part of what determines whether American democracy is governed by careful evaluation of evidence or by whatever narrative is loudest in the moment.
A closing note on humility
Statistical literacy includes humility about one’s own limits. Even readers who have worked through this guide and developed the habits it describes will sometimes be wrong. Statistics are hard; uncertainty is real; even careful readers misread sometimes. The skill is not perfection but improvement on the alternative.
Likewise, polls and statistics are produced by human beings doing their best with imperfect tools in a complex world. The pollsters are not the enemy; the statistical agencies are not the enemy. Most of them are doing serious professional work in good faith. When their results are surprising or uncongenial, the first response should be to read carefully, not to dismiss reflexively. When their results align with one’s prior views, the same careful reading is required. The discipline applies to oneself first, opponents second.
If this guide has been useful, it has been useful in the way good guides usually are: not by replacing the reader’s judgment but by giving them better tools to apply their own judgment. The judgments that follow — about what to believe, what to share, how to vote, how to engage in public life — are the reader’s. The republic continues, in part, by the cumulative effect of those judgments. May they be careful, charitable, and informed.
The bottom line
Statistical literacy is part of the equipment of citizenship in a representative democracy. Polls, economic statistics, crime data, health data, and other measurements are the means by which citizens can know the conditions and views of their fellow Americans — and the means by which government can be held accountable to those conditions. When the measurements break down or are systematically misread, the link between citizen and government strains. Calibrated engagement — trust proportionate to evidence quality, recognition of limits, conviction that careful measurement is possible — is the path. The alternative of credulity or reflexive dismissal is worse for democracy. The ordinary work of careful citizens, multiplied by millions, raises the quality of public discourse; its absence lowers it.
What to read or watch next
- John Dewey, The Public and Its Problems (1927). The classic philosophical statement of what democracy requires of its public, including the conditions for the formation of informed civic judgment.
- James Bryce, The American Commonwealth (1888). The classic outside observer’s account of American democracy; the chapters on public opinion remain illuminating.
- Diane Coyle, Cogs and Monsters: What Economics Is, and What It Should Be (Princeton University Press, 2021). On the broader public-policy stakes of better and worse statistical practice.
- Various works of Hannah Fry and Tim Harford. Both are excellent communicators on the larger civic stakes of statistical literacy; their books and broadcasts are an ongoing resource.
Appendix A: Glossary of Polling and Statistical Terms
This glossary defines terms used throughout this guide. Definitions are written in plain language for general readers; more technical definitions are available in standard statistical references.
AAPOR. American Association for Public Opinion Research, the professional association of public opinion researchers in the United States. Publishes transparency standards, post-election polling reviews, and methodological guidance. Membership in the AAPOR Transparency Initiative (TI) is one signal of methodological seriousness.
Aggregator. A site or service that combines results from multiple polls to produce overall estimates of the state of a race or topic. Major aggregators include FiveThirtyEight, Silver Bulletin, RealClearPolitics, and the New York Times Upshot.
Bias. Systematic error — error in a particular direction — as opposed to random error. A biased poll is consistently off in one direction; an unbiased poll is wrong by random amounts in random directions.
Bradley Effect. The tendency, identified in the 1989 Virginia gubernatorial race, for white respondents to overstate to pollsters their willingness to vote for Black candidates. Now thought to have weakened or disappeared in many contexts.
CPI (Consumer Price Index). The Bureau of Labor Statistics’ measure of inflation, calculated from a fixed basket of goods and services purchased by urban consumers. The most widely cited inflation measure.
Confidence interval. A statistical range constructed so that, if the same poll were repeated many times, a specified percentage of resulting intervals would contain the true population value. Standard published confidence is 95 percent.
Coverage error. Error introduced when a sampling frame does not include all of the target population. The Literary Digest’s 1936 frame had massive coverage error.
Forecast. A prediction of an outcome, often probabilistic. Distinguished from a polling average (which estimates the current state of a race) by the additional uncertainty and information involved in projecting forward.
GDP (Gross Domestic Product). The total value of goods and services produced in an economy in a given period. The standard headline measure of economic output.
House effect. The tendency of a particular pollster to produce results that lean in one direction relative to the average of all pollsters. Aggregators sometimes adjust for house effects.
Likely-voter model. Method for estimating which respondents to an election poll will actually turn out to vote, since polls of all registered voters typically over-represent groups less likely to vote.
Margin of error. Statistical estimate of how much the sample’s results might differ from the true population value due to random sampling variation. Published margins do not capture other sources of error.
Mean / Median / Mode. Three measures of central tendency. Mean is the arithmetic average; median is the middle value; mode is the most common value. They differ for skewed distributions.
NCVS (National Crime Victimization Survey). Bureau of Justice Statistics’ survey of American households about crime victimization. Captures both reported and unreported crime.
NIBRS (National Incident-Based Reporting System). The FBI’s newer crime reporting system, replacing the older Uniform Crime Reporting (UCR) summary system. Captures more detail about each reported crime.
Non-response. Failure of contacted respondents to participate in a poll. Has risen dramatically over the past several decades; non-response bias is the central concern of contemporary polling.
Opt-in panel. An online panel of respondents who have signed up to take surveys, typically for incentives. Different from probability-based panels in that members self-select.
PCE (Personal Consumption Expenditures). Bureau of Economic Analysis measure of inflation, preferred by the Federal Reserve. Differs from CPI in coverage and weighting.
Polling average. Combined estimate from multiple polls, typically weighted by recency, sample size, and sometimes pollster quality.
Population. The full group whose views or characteristics a survey aims to estimate. Examples: American adults, registered Texas voters, likely Republican primary voters.
Probability sample. A sample drawn so that every member of the population has a known, non-zero chance of being included. The gold standard for survey research.
Push poll. A campaign tactic disguised as polling, in which the caller asks loaded questions designed to plant negative information rather than to measure opinion. Widely condemned by professional polling organizations.
Random sampling. Method of drawing a sample in which selection is determined by chance, not by interviewer or respondent preferences. The basis for probability sampling.
Real / Nominal. In economic statistics, real figures are adjusted for inflation; nominal figures are not. Comparisons over time should typically use real figures.
Response rate. Percentage of those contacted by a poll who completed the interview. Has fallen from about 36 percent in the late 1990s to 1–3 percent for many polls today.
Sample. The subset of the population that is actually surveyed. The size and selection method of the sample are central to a poll’s reliability.
Sampling frame. The practical method for reaching members of the target population. Examples: random-digit dialing, address-based sampling, voter registration files, online panels.
Social desirability bias. Tendency of respondents to give answers they consider socially acceptable rather than their honest views. Affects polls on stigmatized topics including racial attitudes, drug use, and politically unpopular candidates.
Sub-sample / subgroup. Portion of the full sample matching some characteristic (“women voters,” “younger voters”). Subgroup results have larger margins of error than the overall sample.
Total survey error. Concept encompassing all sources of error in a survey, including sampling error, coverage error, non-response error, measurement error, and processing error. Typically larger than the published margin of error.
Tracking poll. Poll conducted repeatedly with the same methodology to capture changes over time. Often reports a rolling average.
Turnout. Percentage of eligible voters who actually vote. Varies by election type (presidential, midterm, primary, local) and by demographic group.
U-3 / U-6. Different official measures of unemployment from the Bureau of Labor Statistics. U-3 is the headline rate; U-6 is broader, including discouraged workers and involuntary part-time workers.
Weighting. Statistical adjustment of poll results to make the sample better match the demographic composition of the target population. Standard variables include age, gender, race and ethnicity, education, and region.
Appendix B: Quick-Reference Resources
This appendix gathers the most useful resources mentioned throughout the guide, organized by purpose. Inclusion does not imply endorsement of every position taken by every organization; all are credibly within the mainstream of polling and statistical practice and offer materials of substantial value to careful readers.
Major polling organizations
- Pew Research Center (pewresearch.org). Among the most rigorous and transparent American polling organizations. The Methods 101 series and ongoing methodological writeups are essential reading.
- Gallup (news.gallup.com). Long-running polling organization; publishes the longest continuous series of presidential approval and many other measures.
- Marquette Law School Poll (law.marquette.edu/poll). High-quality academic polling, particularly on Wisconsin and on judicial topics.
- New York Times/Siena College Poll. Among the most rigorous and transparent of media-affiliated polls; full methodology published with each poll.
- Quinnipiac University Poll (poll.qu.edu). Long-running academic-affiliated pollster.
- Marist Poll (maristpoll.marist.edu). Academic-affiliated pollster with strong track record.
- ABC News/Washington Post poll, NBC News poll, Wall Street Journal poll. Major media-affiliated polls with established methodologies and transparent reporting.
Polling aggregators and forecasts
- Silver Bulletin (silverbulletin.com). Nate Silver’s post-FiveThirtyEight project; transparent methodology and probabilistic forecasts.
- FiveThirtyEight (fivethirtyeight.com). Original aggregator brand; readers should check current methodology since the team has changed.
- RealClearPolitics (realclearpolitics.com). Long-running, simpler polling average.
- New York Times Upshot. Polling and political analysis from the Times.
- The Economist forecast. Combines polls and fundamentals; methodology published.
- Decision Desk HQ, Race to the WH, others. Growing field of aggregators with various methodologies.
Federal statistical agencies
- Bureau of Labor Statistics (bls.gov). Unemployment, inflation (CPI), wages, productivity. Methodology publicly documented.
- Bureau of Economic Analysis (bea.gov). GDP, personal consumption expenditures (PCE), trade, regional economic data.
- Census Bureau (census.gov). Population, housing, demographic, business statistics. Authoritative source for many baseline statistics.
- Federal Reserve Economic Data, FRED (fred.stlouisfed.org). Single best free source for economic time series; built by the St. Louis Fed.
- Bureau of Justice Statistics (bjs.ojp.gov). National Crime Victimization Survey and other criminal justice statistics.
- FBI Crime Data Explorer (cde.ucr.cjis.gov). Reported-crime data from law enforcement agencies.
- Centers for Disease Control and Prevention (cdc.gov). Mortality, morbidity, risk factor surveillance.
Public opinion archives
- Roper Center for Public Opinion Research (ropercenter.cornell.edu). The largest archive of public opinion data in the world; maintained at Cornell.
- American National Election Studies, ANES (electionstudies.org). The most rigorous academic study of American voters; conducted around each presidential election.
- General Social Survey, GSS (gss.norc.org). Long-running academic survey of American social attitudes; data freely available for research.
For deeper understanding
- Statistical Modeling, Causal Inference, and Social Science (statmodeling.stat.columbia.edu). Andrew Gelman’s blog; the most consistent source of careful, accessible commentary on polls and statistics.
- AAPOR (aapor.org). The professional association; post-election reviews, transparency standards, resources.
- Our World in Data (ourworldindata.org). Excellent for global comparative data on health, economics, environment, and more.
Recommended introductory books
- Tim Harford, The Data Detective (Riverhead, 2021). Practical rules for citizen statistical literacy.
- Carl Bergstrom and Jevin West, Calling Bullshit (Random House, 2020). Tools for spotting statistical manipulation.
- Charles Wheelan, Naked Statistics (Norton, 2013). Accessible introduction to statistical concepts.
- Darrell Huff, How to Lie with Statistics (Norton, 1954). The classic; still in print, still useful.
- Hannah Ritchie, Not the End of the World (Little, Brown Spark, 2024). Reading health and environmental data.
Appendix C: References and Further Reading
References are organized by chapter, in Chicago Notes and Bibliography style. Citations are intended to enable readers to verify claims and pursue further reading. Many sources have been used across multiple chapters and may appear under more than one chapter heading.
Front matter and Chapter 1
Cantril, Hadley. Gauging Public Opinion. Princeton: Princeton University Press, 1944.
Gallup, George. The Gallup Poll: Public Opinion 1935–1971. 3 vols. New York: Random House, 1972.
Igo, Sarah E. The Averaged American: Surveys, Citizens, and the Making of a Mass Public. Cambridge, MA: Harvard University Press, 2007.
Squire, Peverill. “Why the 1936 Literary Digest Poll Failed.” Public Opinion Quarterly 52, no. 1 (1988): 125–133.
Chapter 2: Sampling
Groves, Robert M., et al. Survey Methodology. 2nd ed. Hoboken, NJ: Wiley, 2009.
Lohr, Sharon L. Sampling: Design and Analysis. 3rd ed. Boca Raton: Chapman and Hall/CRC, 2019.
Pew Research Center. “How Public Polling Has Changed in the 21st Century.” 2023.
Chapter 3: Margin of Error
Manski, Charles F. Public Policy in an Uncertain World. Cambridge, MA: Harvard University Press, 2013.
Manski, Charles F. “Communicating Uncertainty in Official Economic Statistics.” Journal of Economic Literature 53, no. 3 (2015): 631–653.
Chapter 4: Question Wording
Bradburn, Norman M., Seymour Sudman, and Brian Wansink. Asking Questions: The Definitive Guide to Questionnaire Design. San Francisco: Jossey-Bass, 2004.
Schuman, Howard, and Stanley Presser. Questions and Answers in Attitude Surveys: Experiments on Question Form, Wording, and Context. Thousand Oaks: Sage, 1996.
Chapter 5: Response Rates
Keeter, Scott, et al. “What Low Response Rates Mean for Telephone Surveys.” Pew Research Center, 2017.
American Association for Public Opinion Research. “2024 Pre-Election Polling: An Evaluation of the 2024 General Election Polls.” 2025.
Chapter 6: Weighting
Gelman, Andrew, et al. Bayesian Data Analysis. 3rd ed. Boca Raton: CRC Press, 2013.
Wang, Wei, et al. “Forecasting Elections with Non-Representative Polls.” International Journal of Forecasting 31, no. 3 (2015): 980–991.
Chapters 7–9: Election Polls and Aggregators
Erikson, Robert S., and Christopher Wlezien. The Timeline of Presidential Elections: How Campaigns Do (and Do Not) Matter. Chicago: University of Chicago Press, 2012.
Morris, G. Elliott. Strength in Numbers: How Polls Work and Why We Need Them. New York: W.W. Norton, 2022.
Silver, Nate. The Signal and the Noise: Why So Many Predictions Fail — But Some Don’t. New York: Penguin, 2012.
Chapters 10–11: Polling History and Recent Cycles
American Association for Public Opinion Research. Post-election polling reviews of 2008, 2012, 2016, 2020, and 2024. Various dates.
Achen, Christopher H., and Larry M. Bartels. Democracy for Realists: Why Elections Do Not Produce Responsive Government. Princeton: Princeton University Press, 2016.
Chapter 12: How to Lie with Statistics
Huff, Darrell. How to Lie with Statistics. New York: W.W. Norton, 1954.
Tufte, Edward R. The Visual Display of Quantitative Information. 2nd ed. Cheshire, CT: Graphics Press, 2001.
Wheelan, Charles. Naked Statistics: Stripping the Dread from the Data. New York: W.W. Norton, 2013.
Chapter 13: Issue Polls
Converse, Philip E. “The Nature of Belief Systems in Mass Publics.” In Ideology and Discontent, edited by David E. Apter, 206–261. New York: Free Press, 1964.
Zaller, John R. The Nature and Origins of Mass Opinion. Cambridge: Cambridge University Press, 1992.
Chapter 14: Economic Statistics
Coyle, Diane. GDP: A Brief but Affectionate History. Princeton: Princeton University Press, 2014.
Bureau of Labor Statistics. “Handbook of Methods.” Washington, DC: U.S. Department of Labor. Updated periodically.
Chapter 15: Crime, Health, and Other Counts
Bureau of Justice Statistics. “Criminal Victimization” series. Annual.
Ritchie, Hannah. Not the End of the World. New York: Little, Brown Spark, 2024.
Chapters 16–18: The Citizen’s Stake
Bergstrom, Carl T., and Jevin D. West. Calling Bullshit: The Art of Skepticism in a Data-Driven World. New York: Random House, 2020.
Dewey, John. The Public and Its Problems. 1927. Reprint, Athens, OH: Swallow Press, 1991.
Harford, Tim. The Data Detective: Ten Easy Rules to Make Sense of Statistics. New York: Riverhead, 2021.
Lippmann, Walter. Public Opinion. New York: Harcourt, Brace, 1922.
Herbst, Susan. Numbered Voices: How Opinion Polling Has Shaped American Politics. Chicago: University of Chicago Press, 1993.
A note on additional sources
This guide has drawn on a much larger pool of materials than any reference list could fully document, including ongoing reports from the polling organizations, statistical agencies, and aggregators listed in Appendix B. Readers seeking deeper engagement with any particular question are encouraged to follow the citations in the works above; the field is well-served by serious scholars and serious practitioners. The conversation between citizens and the data that describes their world is ongoing, and every careful reader of this guide is now part of it.
Ask this guide
Ask a question and get an answer drawn only from Polls and Statistics. It won't make things up — if the answer isn't here, it'll tell you.
