Born Between 2Generals

Join the Legacy
All Fact Guides

Media & Information

Reading Research

How to read a study and judge what it really shows.

132 min read · 29,041 words

“It ain’t what you don’t know that gets you into trouble. It’s what you know for sure that just ain’t so.”

Foreword

On August 28, 2015, the journal Science published the results of an unusual experiment. A team of 270 researchers — the Open Science Collaboration, led by University of Virginia psychologist Brian Nosek — had spent four years attempting to replicate 100 published findings from three prestigious psychology journals. The original studies had found statistically significant effects in 97 of 100 cases. The replications, conducted with high statistical power and largely the same methods, found significant effects in only 36. On average, the effect sizes of the replications were about half what the original studies reported.

The reaction was substantial. Some commentators concluded that psychology was in crisis; others argued that the replication failures were evidence the field was working as it should, identifying and correcting errors. Subsequent replication projects in cancer biology, experimental economics, social psychology, and other fields produced similar patterns of partial replication. By the late 2010s, what had started as a methodological dispute within psychology had become a broader conversation, sometimes called “the replication crisis,” about how confident we should be in any single published research finding.

The crisis prompted serious reform: preregistration of studies, open data and code, registered reports, larger sample sizes, more careful statistical practice. The reforms have made some areas of science more rigorous than they were a decade ago. They have not made science easier for the citizen reader; if anything, the lesson of the past decade has been that even competent professionals routinely produce findings that turn out not to be reliable. The naive position — “it’s in a peer-reviewed journal, so it must be true” — has been retired by people who study research practice. The cynical position — “all studies are unreliable, so I will trust whichever ones agree with me” — is worse, but it is the natural alternative if the naive position fails.

This guide is for the citizen who wants to do better than either: who wants to read research findings carefully, weight them appropriately, recognize the difference between strong and weak evidence, and understand where in any given field the science is settled and where it remains in flux. The work is harder than the naive position made it look, but it is not as hard as the cynical position implies. Most of the time, with a few practical tools, a careful citizen can form a reasonable view of how confident to be in a particular research finding. This guide tries to provide those tools.

A note on what this guide is and is not. It is not a substitute for becoming a scientist, or even for taking an introductory statistics course. It is a practical guide for the citizen reader: someone who encounters research findings in news stories, in policy debates, in their own decisions about health and life, and who wants to evaluate those findings with appropriate skill. The chapters move from the basic structure of scientific research, through the major problems researchers and reformers are working on, into the practical work of reading specific studies and understanding when scientific consensus deserves deference.

Three framing principles. First, science is real and works. The methodological problems discussed in this guide are real, and they affect particular findings, but they do not mean that science as a whole is broken. The same scientific community that has failed to replicate certain findings has also produced extraordinary advances in physics, biology, medicine, and many other fields. The replication crisis is most acute in some social and biomedical sciences; it does not extend evenly across all of science. Second, peer review is a useful but imperfect filter. Peer-reviewed research is, on average, more reliable than non-peer-reviewed research; that does not make it certainly correct. Third, scientific consensus deserves substantial weight without being treated as infallible. Where genuine scientific consensus exists — evolution, vaccine safety, anthropogenic climate change — it represents the considered judgment of large numbers of working scientists across many institutions, and citizens should give such consensus great weight. Where consensus is absent or contested, citizens should be aware of the genuine uncertainty rather than seizing on whichever side of the dispute matches their priors.

This is a guide to reading research as a citizen, not as a scientist. The standard a citizen needs is not the standard a peer reviewer needs. Citizens need to know when a finding is reliable enough to act on, when it deserves further investigation, when it is too preliminary to draw conclusions from, and when the broader body of evidence on a topic is or is not settled. These are practical judgments that careful reading can support. They are the kinds of judgments self-government in a scientifically informed era requires.

PART ONE

How Science Works

The basic structure: what a scientific study actually is, what peer review does, and how to read a research paper

CHAPTER 1

What a Scientific Study Actually Is

A scientific study is a specific, structured attempt to learn something about the world by collecting and analyzing evidence in a way that is documented well enough that other people can examine, criticize, and try to replicate the work. The specific designs vary enormously across fields — a particle physics experiment, a randomized clinical trial, a sociological survey, a paleontological dig, an econometric analysis of historical data — but the underlying logic is shared. A claim about the world is made; evidence is gathered that bears on the claim; the analysis of that evidence is documented in enough detail that the work can be checked. The reliability of any particular finding rests not on the authority of the researchers but on the quality of the evidence and the soundness of the analysis.

The basic structure

Most scientific studies, regardless of field, share certain elements:

  • A research question. The specific question the study seeks to answer. “Does drug X reduce mortality in patients with disease Y?” “Does increasing the minimum wage reduce employment?” “Does microblogging by political leaders affect polarization?” Good research questions are specific enough that they can be answered with the available evidence; vague questions tend to produce vague answers.
  • A hypothesis. The specific prediction the study tests. Often a hypothesis takes the form: “If we do A, we will see B.” Hypotheses can be either confirmatory (testing predictions made in advance based on prior theory) or exploratory (looking for patterns in the data without fixed predictions). The distinction matters: exploratory findings should be treated as hypotheses for future testing, not as confirmed conclusions.
  • A study design. The plan for collecting evidence: who or what will be studied, how the data will be collected, what comparisons will be made, what controls will be used. Different designs have different strengths and weaknesses; Chapter 4 discusses these in detail.
  • Data collection. The actual gathering of evidence. The methods used, the population sampled, the measurements taken, and the procedures followed are all part of the study’s record.
  • Analysis. The statistical or qualitative methods used to evaluate the evidence. Modern published research typically uses statistical analyses appropriate to the data; the specific tests, the assumptions they require, and the way results are interpreted are part of what reviewers and replicators evaluate.
  • Conclusions. What the study’s authors believe the evidence supports, and with what level of confidence. Honest conclusions acknowledge the limits of the study and the alternative interpretations that remain possible.
  • Documentation and dissemination. Publication, typically in a peer-reviewed journal, in enough detail that other researchers can evaluate, replicate, or extend the work. Increasingly, raw data and analysis code are also published, allowing direct verification of the analyses.

The role of falsifiability

A useful philosophical principle, articulated by Karl Popper, is that scientific claims should be falsifiable: it should be possible to specify in advance what evidence would count against the claim, not just what evidence would support it. A claim that no possible observation could refute is not really a scientific claim; it is doing some other kind of work. The principle is more nuanced than it appears (almost any claim can be defended against any observation by adjusting auxiliary assumptions), but it captures something important: good scientific claims make commitments that could be wrong, and the willingness to make such commitments is part of the discipline.

Citizens reading research can apply a related test: what would the researchers acknowledge as evidence that their conclusion was wrong? In many honest research papers, the authors explicitly discuss what alternative interpretations remain consistent with the data and what additional studies would help discriminate among them. In less honest work, the author treats their preferred interpretation as the only one supported by the evidence and dismisses alternatives. The willingness to engage genuinely with alternative interpretations is one mark of intellectually serious research.

Theory, observation, and the dance between them

Science is sometimes presented as if it consisted purely of unbiased observation, with theories simply emerging from the data. The reality is more complex. Researchers approach data with theoretical commitments and prior expectations; the choice of what to study, what to measure, what counts as a meaningful pattern, all depend on theoretical frameworks. This is not a flaw of science; it is a feature. Theories tell us where to look, what to compare, and how to interpret what we find. The discipline is in being willing to revise theories when the evidence demands it, not in pretending to have no theoretical commitments.

This means that the same data can sometimes support multiple theoretical interpretations. The work of science is partly the work of choosing among them based on which is best supported, which makes the most accurate predictions, which integrates best with other accepted theories, which is most parsimonious. None of these criteria is mechanical; reasonable scientists can disagree. The image of science as a single body of agreed-upon facts is a simplification; in any field at the research frontier, there is genuine ongoing dispute about how to interpret the evidence.

What kinds of questions science can address

Science is exceptionally good at answering empirical questions: questions about what is, what was, what causes what, how things work. Science is less suited to normative questions about what should be, what is good, what is just. The two often appear together — an empirical claim about the consequences of a policy is bound up with normative claims about what consequences are desirable — but the methods of science are most reliable on the empirical part.

This distinction matters for citizen reading. A research paper claiming to identify the consequences of a policy is doing empirical work; a research paper claiming the policy should be adopted is doing normative work that goes beyond what the empirical evidence alone can support. Both are legitimate but they are different. Conflating them — either by reading empirical findings as normative recommendations or by dismissing empirical findings because the normative implications are unwelcome — leads to bad reasoning.

The bottom line

A scientific study is a structured attempt to learn about the world by collecting evidence in a way that can be examined and replicated. Studies share basic elements (question, hypothesis, design, data, analysis, conclusions, documentation) regardless of field. Falsifiability — the willingness to specify what would count against the claim — is a useful test. Theory and observation are intertwined; data alone do not interpret themselves. Science is best at empirical questions; normative questions go beyond what science alone can settle. Reading research carefully means engaging with each of these elements rather than relying on the authority of the source.

What to read or watch next

  • Karl Popper, The Logic of Scientific Discovery (1934, English 1959). The classic statement of falsifiability and the philosophy of science.
  • Thomas Kuhn, The Structure of Scientific Revolutions (University of Chicago Press, 1962). The complementary view that science proceeds through paradigm shifts shaped by theoretical commitments. Influential and contested.
  • Stuart Ritchie, Science Fictions: How Fraud, Bias, Negligence, and Hype Undermine the Search for Truth (Metropolitan Books, 2020). The most useful single book on what goes wrong in contemporary science and what to do about it. Excellent starting point.

CHAPTER 2

Peer Review and What It Does

Most published academic research has gone through peer review: a process in which the authors’ manuscript is sent by a journal editor to other researchers in the field, who read it and recommend whether to publish, revise, or reject. The reviewers are typically anonymous to the authors; their identity is known only to the editor. The process is a central feature of how scientific publishing works and is widely cited as a key reason peer-reviewed research is more reliable than non-peer-reviewed work. It is also widely misunderstood. Peer review is a useful filter, but it is much weaker than its reputation suggests. Understanding what it actually does is essential for reading research carefully.

What peer review is supposed to do

In the standard model, peer review serves several functions:

  • Quality control. Reviewers check that the methodology is sound, the analysis is appropriate, the claims match what the data show, and the writing is clear. Manuscripts with serious flaws are revised or rejected.
  • Filtering. Journals receive far more submissions than they can publish. Peer review helps editors decide which manuscripts merit publication in their journal.
  • Improvement. Reviewers often suggest improvements: additional analyses, clearer presentation, important alternative interpretations to discuss. Many published papers are substantially better than the original submissions because of reviewer feedback.
  • Gatekeeping. Peer review serves to maintain the integrity of the field by keeping clearly bad work out of the literature.

What peer review actually does

Empirical research on peer review has consistently found that it does these things less reliably than the standard story suggests. Several patterns are well-documented:

  • Reviewers rarely catch errors. When researchers have deliberately introduced errors into manuscripts and sent them to reviewers, reviewers typically catch a small fraction of the errors. Statistical errors are particularly likely to slip through; many reviewers do not have the statistical expertise to evaluate the analyses they are reviewing, and many published papers contain statistical mistakes that reviewers did not flag.
  • Inter-reviewer agreement is low. The same manuscript sent to two reviewers often produces sharply different recommendations. The decisions are noisier than the standard story suggests.
  • Major frauds are usually detected after publication, not in review. The most prominent cases of scientific misconduct in recent years (Diederik Stapel’s fabricated data in social psychology, Marc Hauser’s problems at Harvard, the cases in stem cell research, and many others) were typically identified by replication failures, statistical anomalies in the published papers, or whistleblowers — not by reviewers during peer review.
  • Bias in review is real. Studies have documented patterns by which reviewers favor manuscripts from prestigious institutions, manuscripts whose findings align with their prior views, and manuscripts whose authors they recognize. Double-blind review (in which authors are also anonymous to reviewers) reduces some of these biases but is not universally used.
  • Predatory journals exist. A growing number of journals charge fees for publication and provide little or no peer review, lending the appearance of peer-reviewed legitimacy to work that has not been seriously evaluated. Distinguishing reputable journals from predatory ones is a non-trivial skill.

How to use peer review as a signal

Despite these limits, peer review is still a useful signal. Peer-reviewed research is, on average, more reliable than non-peer-reviewed research, just much less reliably than the standard story suggests. A few practical principles:

  • Treat peer review as a low filter, not a high one. Publication in a peer-reviewed journal means the work passed a minimum bar; it does not mean the work is necessarily correct.
  • Pay attention to journal reputation. Some journals (Nature, Science, Cell, the New England Journal of Medicine, the Journal of the American Medical Association, the Quarterly Journal of Economics, and others) are highly selective and apply rigorous review. Others publish almost anything submitted. The journal’s reputation is a rough proxy for the rigor of its review.
  • Be wary of predatory journals. The Beall’s List archive (web archive, since the original list was taken down) and similar resources document journals with predatory practices. Publication in such a journal does not constitute meaningful peer review.
  • Look for replication and citation. Findings that have been replicated and built on by other researchers are more reliable than findings that exist only in their original publication.
  • Note the time since publication. Recent findings, even in major journals, are more provisional than findings that have stood up to a decade of scrutiny. The replication failures of the past decade have been concentrated in newer findings; well-established findings have generally held up better.
  • Read post-publication review. For many findings, post-publication review (commentaries, replies, replications, citations in subsequent work) provides better evaluation than the initial peer review. Sites like PubPeer host substantive critiques of published papers; Google Scholar shows citation patterns and links to follow-up work.

Preprints and the changing publishing landscape

In some fields, particularly physics, mathematics, and increasingly biology and economics, researchers post their papers as preprints (on arXiv, bioRxiv, medRxiv, SSRN, NBER, and similar servers) before or instead of formal peer review. This has accelerated communication of research findings but raises its own questions about evaluation.

Preprints are not peer-reviewed; their reliability depends entirely on the work itself, the reputation of the authors, and post-publication review. Some preprints will eventually be published in peer-reviewed venues; others will not. Citizens encountering a preprint claim should treat it as preliminary evidence — worth attention, but lower confidence than peer-reviewed work in a reputable journal. The COVID-19 pandemic was a particularly visible test of this: many important findings appeared first as preprints, sometimes accurately, sometimes not, and journalists and policymakers had to develop the habit of checking whether preprints had been peer-reviewed before treating them as established.

The bottom line

Peer review is a useful but weaker filter than its reputation suggests. It rarely catches errors, has substantial inter-reviewer disagreement, and rarely detects major fraud. Peer-reviewed research is, on average, more reliable than non-peer-reviewed research, but publication in a peer-reviewed journal is a low bar rather than a high one. Journal reputation, replication, citation patterns, post-publication critique, and time-tested findings are better signals of reliability than peer review status alone. The growing role of preprints adds another layer that requires the citizen reader to check whether a claim has actually been peer-reviewed.

What to read or watch next

  • Stuart Ritchie, Science Fictions (Metropolitan Books, 2020). Particularly the chapter on peer review and how it actually functions.
  • Richard Smith, “Peer Review: A Flawed Process at the Heart of Science and Journals,” Journal of the Royal Society of Medicine 99 (2006): 178–182. The former British Medical Journal editor’s critical assessment, by someone who ran a major journal for many years.
  • PubPeer (pubpeer.com). Platform for post-publication peer review; useful for checking whether published papers have been substantively criticized.

CHAPTER 3

Anatomy of a Research Paper

A typical research paper in most empirical fields has a recognizable structure. Once you know the structure, you can navigate any paper in the field reasonably efficiently, finding the parts that matter for your purposes without reading every word. This chapter walks through the standard structure and what each section is for, with practical advice on what to read carefully and what to skim.

The standard structure

Most empirical research papers follow the IMRaD format: Introduction, Methods, Results, and Discussion. The structure may have slightly different names in different fields (Background, Materials, Findings, Conclusion), but the underlying pattern is consistent. There is also typically an Abstract at the beginning and a References list at the end.

Abstract

A short summary of the paper, typically 150–300 words, usually structured like a miniature paper itself: background, methods, results, conclusion. The abstract is the most-read part of any paper. For most readers, the abstract plus the figures plus the discussion section is enough to evaluate whether the paper is relevant and what it claims to show.

A few warnings about abstracts: they are often more confident than the body of the paper warrants. Authors compress nuance to fit the word limit and to maximize impact. Important caveats and limitations may appear only in the discussion or methods sections. Reading only abstracts produces a systematically more confident impression of the literature than the literature actually warrants.

Introduction

The introduction lays out the research question, why it matters, and what previous research has shown. A good introduction explains the gap in existing knowledge that the current study aims to fill, presents the specific hypotheses or research questions, and gives the reader enough context to evaluate the importance of the contribution.

For citizen readers, the introduction is useful for orientation: what question is being asked, why this question, what has been studied before. Less useful: the framing of why this study is novel and important, which can be exaggerated for publication purposes.

Methods

The methods section describes how the study was conducted: who or what was studied, how the data were collected, what statistical analyses were used, what software and procedures were employed. In rigorous fields, methods sections are detailed enough to allow another researcher to attempt replication.

This is the most important section for evaluating reliability. The strengths and weaknesses of the study are typically more visible in the methods than anywhere else. Was the sample size adequate? Were appropriate controls included? Were the measurements reliable? Were the analyses appropriate to the data? A reader who skims the methods to get to the results is reading the paper backward; the methods are what make the results meaningful.

Results

The results section presents the empirical findings, typically with tables, figures, and statistical tests. In careful papers, the results are reported as fully as possible, including findings that did not support the authors’ hypotheses. In less careful papers, results that did not support the preferred conclusion may be downplayed or relegated to supplementary material.

Reading results well requires some statistical literacy. The numbers reported (effect sizes, confidence intervals, p-values) all have specific meanings; treating them as authoritative without understanding what they say is a common error. Chapters 5 and 6 discuss this in detail.

Discussion

The discussion interprets the results in light of prior research, considers alternative explanations, acknowledges limitations, and articulates implications for the field and possibly for policy or practice. A strong discussion engages seriously with limitations and alternative interpretations; a weak discussion glosses over them in favor of a triumphal narrative.

For citizen readers, the discussion is often where to look for the careful interpretation. The abstract may overstate; the results section may be dense; the discussion is typically where the authors lay out what they think the findings mean, with what level of confidence. Reading the discussion carefully often reveals important caveats that headline coverage misses.

References

The list of cited works. Useful for finding related research, evaluating which prior work the authors are engaging with, and identifying the key references in a field. An unusually thin reference list (few citations to existing work) or an unusually self-citation-heavy list (many references to the authors’ own previous work) can be signals about how the paper relates to the broader field.

Supplementary materials

Most modern papers have supplementary materials available online: additional analyses, robustness checks, raw data, code, materials. Increasingly, these are an essential part of the paper rather than a footnote. A reader who wants to evaluate the rigor of a study should at minimum check whether supplementary materials are available and what they contain. Open data and code are particularly valuable for verification.

How to read a paper as a citizen

Reading a research paper from beginning to end is rarely the most efficient strategy. A more useful approach for most citizen readers:

  • Start with the abstract. Get the basic claim. Note any ways the abstract seems to overstate.
  • Read the figures and figure captions. Visual presentation often communicates the key findings more clearly than the text. Reading the figures with care often reveals the strengths and weaknesses of the study faster than reading the prose.
  • Skim the methods for major issues. Sample size, study design, key procedures. Major flaws often jump out from quick skimming; details can be revisited if needed.
  • Read the discussion section. Particularly the limitations and alternative interpretations. This is where careful authors flag what their study cannot show.
  • Check the citations. What does this paper cite, and is it cited? Tools like Google Scholar make this quick. A paper rarely cited may be a niche contribution; a paper widely cited may be a more central reference.
  • Consider the source. Author affiliations, funding sources, and the journal’s reputation are all signals. None is definitive, but together they help calibrate how much weight to put on the findings.

The bottom line

Empirical research papers follow a standard structure (Abstract, Introduction, Methods, Results, Discussion, References, Supplementary). Each section serves a specific purpose. Abstracts often overstate; methods reveal the study’s strengths and weaknesses; results require statistical literacy to read well; discussions are where careful authors flag caveats. For citizen readers, an efficient strategy is to start with the abstract, read the figures, skim the methods for issues, read the discussion carefully, and consider citation patterns and sources. Reading carefully takes practice but is learnable, and the same skills apply across fields.

What to read or watch next

  • Cailin O’Connor and James Owen Weatherall, The Misinformation Age: How False Beliefs Spread (Yale University Press, 2019). On how scientific findings spread (and mis-spread) through scientific and public networks.
  • John Bohannon, “Who’s Afraid of Peer Review?” Science 342 (2013): 60–65. The famous sting operation in which 157 of 304 open-access journals accepted a fake paper with obvious flaws. Sobering for any reader.
  • Retraction Watch (retractionwatch.com). Excellent ongoing journalism on retractions, misconduct, and the integrity of the scientific literature. Useful for spotting problems in specific cases.

PART TWO

Evaluating a Single Study

Designs and what they can show, sample size and effect size, and what p-values and confidence intervals actually mean

CHAPTER 4

Study Designs and What They Can Show

Not all studies are created equal. Two papers can each be carefully done, peer-reviewed, and published in respectable journals, and yet the conclusions you should draw from them can differ enormously — because the underlying study designs differ. The single most important thing a citizen reader can learn is how to recognize the design of a study and what kinds of conclusions that design can support. A randomized clinical trial of a new drug speaks with much more authority about whether the drug works than an observational study following people who happened to take the drug. A meta-analysis combining dozens of similar studies tells you something different from any single one of them. A case study of a single unusual patient tells you something different again. Reading the design — not just the headline — is the basic act of statistical literacy.

The hierarchy of evidence

In medicine and increasingly in other fields, study designs are often arranged in a rough hierarchy from weakest to strongest. The hierarchy is not absolute — a well-done observational study can sometimes be more useful than a poorly done randomized trial — but as a rule of thumb it is genuinely useful. From bottom to top, the hierarchy looks roughly like this:

  • Anecdote and expert opinion. Stories about individual cases or the considered judgment of experienced practitioners. Sometimes the only available evidence; not by itself a strong basis for general claims.
  • Case reports and case series. Descriptions of individual patients or small groups, often the first signal that something interesting is happening. Cannot establish frequency or causation but can generate hypotheses.
  • Cross-sectional studies. Snapshots of a population at one point in time, measuring exposures and outcomes simultaneously. Useful for prevalence, weak for causation.
  • Case-control studies. Compare people who have an outcome (cases) to people who do not (controls), looking backward at exposures. Efficient for rare outcomes; vulnerable to recall bias and selection bias.
  • Cohort studies. Follow groups forward in time, observing who develops the outcome. More reliable than case-control but expensive; results can take years.
  • Randomized controlled trials. Subjects are assigned at random to receive the intervention or a control. Considered the gold standard for causation in medicine and increasingly elsewhere.
  • Systematic reviews and meta-analyses. Combine the results of multiple prior studies under explicit criteria. The best ones synthesize the evidence on a question; the worst inherit the biases of the studies they include.

Experimental versus observational

The most important distinction within the hierarchy is between experimental designs and observational designs. In an experimental design, the researcher controls who receives the treatment. In an observational design, the researcher records what is already happening. The difference matters because experimental designs, if randomization is done well, eliminate most sources of confounding — the lurking variables that can make a non-causal correlation look like a causal one. Observational designs, no matter how large or carefully analyzed, can never fully eliminate confounding. They can only adjust for the confounders the researchers thought to measure.

This is why, in fields where randomized experiments are possible, they tend to be considered the strongest design. A drug company demonstrating that their pill reduces heart attacks compared to placebo, in a randomized trial of 20,000 patients, is making a much stronger claim than an observational study showing that people who happen to take the pill have fewer heart attacks. The observational result might reflect the drug; it might also reflect that people who get prescribed the pill are different in some other way — wealthier, more health-conscious, less sick to begin with — from those who do not.

When randomization is impossible

Many of the questions citizens care about cannot be answered by randomized experiments. We cannot randomly assign children to grow up in two-parent or one-parent households; we cannot randomly assign people to live near major highways or far from them; we cannot randomly assign nations to have different political systems. For these questions, observational research is the only option, and the methodological challenge is enormous. Modern social science has developed sophisticated techniques to approximate causal inference from observational data — instrumental variables, regression discontinuity, difference-in-differences, propensity score matching, and natural experiments — but none is as clean as a randomized trial. Citizens reading research in these areas should be aware that even excellent observational studies are typically more uncertain than excellent experimental ones.

Natural experiments

A natural experiment is a situation in which something close to randomization happens by accident. A draft lottery randomly assigns young men to military service; a court system randomly assigns judges to cases, some judges more lenient than others; a state policy changes at a particular date but not in a neighboring state. In each case, the haphazardness of the situation provides leverage for causal inference that pure observation does not. Some of the most influential social science of the past forty years — work by economists like David Card, Joshua Angrist, and Guido Imbens, recognized with Nobel Prizes — has used natural experiments to answer questions that randomized experiments cannot ethically address.

Specific designs in detail

Case reports and case series

A case report describes one patient or a small handful in detail. The first reports of AIDS, in 1981, were case series. The first reports of thalidomide-induced birth defects were case series. Case reports cannot establish how common a phenomenon is or what causes it, but they are often the first signal that something is happening. The citizen reading a case report should treat it as a hypothesis-generator, not a finding. The fact that a single doctor saw two unusual cases does not mean a new disease has been identified, but it might mean it has.

Cross-sectional studies

A cross-sectional study measures both an exposure and an outcome at the same time, in a population of people. They are common in epidemiology and survey research. Their fundamental weakness is that they cannot tell you whether the exposure caused the outcome or the outcome caused the exposure or both reflect something else. A finding that depressed people exercise less can mean that depression reduces exercise, that lack of exercise causes depression, or that some third factor (chronic illness, poverty, social isolation) causes both. Cross-sectional designs are useful for prevalence — “What fraction of the population has condition X?” — but weak for causation.

Case-control studies

A case-control study identifies a group of people who already have the outcome of interest (cases) and compares them to a group who does not (controls), looking backward to see what exposures differed between them. The classic case-control design was used to identify the link between smoking and lung cancer in the early 1950s. Case-control studies are efficient for rare outcomes — you do not have to follow tens of thousands of people forward to find the few who develop the condition — but they suffer from two notorious problems. Recall bias: cases asked about past exposures often remember more accurately than controls, because they have been thinking about possible causes. Selection bias: the way controls are chosen can introduce systematic differences from cases unrelated to the exposure of interest.

Cohort studies

A cohort study identifies a group of people, characterizes their exposures or characteristics, and follows them forward in time to see what outcomes develop. Cohort studies are more expensive than case-control studies and can take years or decades — the Framingham Heart Study, started in 1948, has now followed three generations of residents of one Massachusetts town. They are considered the strongest observational design because they avoid recall bias (exposures are recorded before outcomes occur) and most forms of selection bias. They cannot, however, control for unmeasured confounders. A cohort study showing that people who eat a Mediterranean diet have less heart disease cannot tell you whether the diet itself helps, or whether people who eat that way also exercise more, smoke less, or have higher incomes — unless those factors are measured and controlled for in the analysis.

Randomized controlled trials

In a randomized controlled trial (RCT), researchers randomly assign participants to receive either the intervention or a control. The randomization, if properly done, ensures that on average the two groups are similar in every respect except the treatment. Differences in outcome can therefore be attributed to the treatment with high confidence. RCTs are the gold standard for causal inference but they are expensive, sometimes ethically impossible, and have their own failure modes. Subjects can drop out unevenly between groups; blinding (concealing which group is which from participants and researchers) can fail; the population enrolled in the trial may differ from the population to which results are later generalized; effect sizes seen in tightly controlled trials sometimes shrink when treatments are deployed in messy real-world settings.

Systematic reviews and meta-analyses

A systematic review identifies all studies that have addressed a particular question, evaluates their quality, and summarizes their findings using explicit criteria. A meta-analysis goes further and combines the results statistically, producing a single estimate of the effect with a confidence interval. Done well, these are among the strongest forms of evidence available, because they pool data across many studies and reduce the role of chance. Done poorly, they inherit and amplify the biases of the studies they include — if all the included studies are flawed in the same way, the meta-analysis will be confidently wrong. The Cochrane Collaboration, an international network founded in 1993, produces particularly rigorous systematic reviews in medicine and is a useful starting point for citizens trying to understand the state of evidence on a clinical question.

Qualitative research

Not all research is quantitative. In fields like sociology, anthropology, education, and parts of public health, qualitative methods — in-depth interviews, ethnographic observation, focus groups, content analysis — produce rich descriptions of social phenomena that quantitative methods cannot. Qualitative research has a different logic. It is not designed to generalize to populations from samples; it is designed to develop nuanced understanding of how things work in specific contexts. A citizen reading a qualitative study should not ask “How many subjects, statistically significant?” They should ask “Does the description ring true? Does the analysis fit the evidence presented? Are alternative interpretations considered?” Both quantitative and qualitative work can be done well or badly; conflating their standards — dismissing qualitative work for not being statistical, or dismissing quantitative work for being reductive — reflects unfamiliarity with the different things they are trying to do.

The bottom line

Study designs differ enormously in what they can support. Experimental designs (especially RCTs) are best for causation; observational designs are necessary when experiments are impossible but require careful interpretation. Case reports generate hypotheses; cross-sectional studies estimate prevalence; case-control and cohort studies trade efficiency against bias control; meta-analyses pool evidence across studies. The first question a careful reader asks is: what is the design, and what kind of conclusion can it actually support?

What to read or watch next

  • Trisha Greenhalgh, How to Read a Paper: The Basics of Evidence-Based Medicine (Wiley-Blackwell, 6th ed., 2019). The standard introduction for medical professionals; chapters 3–7 walk through the major designs in plain English.
  • Joshua Angrist and Jörn-Steffen Pischke, Mostly Harmless Econometrics (Princeton University Press, 2009). The accessible introduction to natural experiments and causal inference in observational data. Technical but unusually clear.
  • Cochrane Collaboration (cochrane.org). The leading source of systematic reviews in medicine; the plain-language summaries are a model for how complex evidence can be communicated to non-specialists.
  • Esther Duflo, Abhijit Banerjee, and Michael Kremer, Nobel lectures (2019). The trio describe how randomized experiments transformed development economics; available at nobelprize.org.

CHAPTER 5

Sample Size, Effect Size, and Statistical Power

If you read only one chapter of this guide, this is the one. Most of the times research findings turn out to be wrong, the immediate technical reason has something to do with the relationships among three quantities: how many subjects were studied, how big the effect being claimed is, and how reliably the study could have detected an effect of that size if one really existed. The relationships among these three quantities go by the technical name of statistical power, and the failures of statistical power are the single most consistent thread running through the replication crisis. A citizen who internalizes the basic logic of power can spot, without doing any math at all, why a great many published findings should be regarded skeptically.

Sample size

Sample size is simply the number of subjects in the study. Other things equal, larger samples produce more reliable estimates than smaller ones. A study of 30 people is much noisier than a study of 3,000. The reason is that any sample includes random variation — some subjects respond differently than others for reasons unrelated to the treatment — and larger samples average out the noise. With 30 subjects, a few unusual responders can swing the apparent result substantially; with 3,000, they cannot. Citizens reading research should always note the sample size and ask whether it is plausibly large enough to support the claim being made. A study claiming to demonstrate a small effect with a small sample should be regarded with particular suspicion.

Effect size

Effect size is how big the difference is between groups, or how strong the relationship is between variables. Effect size is reported in many different units depending on the field: a difference in means measured in standard deviations (Cohen’s d), a correlation coefficient (r), an odds ratio, a relative risk, a percentage point difference. The specifics matter less than the underlying point: an effect can be statistically significant — unlikely to be due to chance — and yet too small to matter in practice. A drug that lowers blood pressure by 0.2 millimeters of mercury, in a study large enough to detect that tiny effect, will reach statistical significance without offering any real benefit to patients. A study of educational interventions can show “statistically significant” gains that translate into one extra correct answer per hundred questions. Statistical significance is a yes-or-no determination; effect size is the magnitude that determines whether the result has any practical importance.

This is why one of the most consistent forms of misreading is to attend to the asterisks and ignore the numbers. A finding marked p < 0.05 with great fanfare may correspond to an effect that, if real, is far too small to act on. Conversely, a study showing a large effect that just barely fails to reach significance — perhaps because the sample was small — may be telling you something important that further research will confirm. Reading effect sizes alongside significance is essential.

Statistical power

Statistical power is the probability that a study will detect an effect, if an effect of a given size really exists. Power depends on sample size, effect size, and the variability of the measurements. A study has high power if it is large, the effect is big, and the noise is low. A study has low power if it is small, the effect is subtle, or the measurements are messy. Most published research in fields touched by the replication crisis has been chronically underpowered: studies with sample sizes of 30 or 50 or 100 attempting to detect modest effects, with the result that even when effects are real, the studies miss them about half the time.

A subtle implication is that low-powered studies that do find significant effects are particularly suspect. If a study has, say, 30 percent power to detect an effect of a given size, it will miss that effect 70 percent of the time. The 30 percent of trials that find the effect will, on average, exaggerate its size, because the only way an underpowered study can produce a significant result is by getting lucky with a sample that happens to show a larger-than-true effect. This phenomenon, known as the winner’s curse or effect-size inflation, is one of the central reasons replication studies tend to find smaller effects than the originals — even when the underlying effect is real.

How big is big enough?

There is no universal answer to “how big should the sample be?” — it depends on the effect size you care about and the variability of your measurements. As a rough orienting guide, however, a few benchmarks may be useful. To detect a moderate-sized difference between two groups (Cohen’s d of about 0.5) with 80 percent power and a conventional significance threshold, you need roughly 64 subjects per group, or 128 total. To detect a small effect (d of 0.2), you need closer to 400 per group, or 800 total. Many of the published psychology studies that failed to replicate in the 2010s had sample sizes of 30 to 60 total. They were chronically underpowered to detect anything but enormous effects, yet routinely reported finding small ones, because of the selection process by which only the lucky-sample versions got published.

Confidence intervals

Closely related to sample size and effect size is the confidence interval. A confidence interval is a range of values around an estimate, calculated so that, in repeated samples, the true value would fall inside that range a specified percentage of the time — conventionally 95 percent. A study estimating that a drug reduces mortality by 10 percent might report a 95 percent confidence interval of 4 to 16 percent. The interpretation is that the estimate is 10 percent, but the data are also consistent with effects ranging from 4 percent to 16 percent. Wide confidence intervals — a study finding 10 percent reduction with confidence interval of −2 to 22 percent — indicate that the data have not pinned down the answer very precisely. They are common in small studies.

Confidence intervals are more informative than p-values because they convey both the size of the estimated effect and the precision of the estimate. A finding of “10 percent reduction (95 percent confidence interval: 4 to 16)” is more useful than “p = 0.003.” Modern style guides in many fields explicitly recommend reporting confidence intervals alongside or instead of p-values for this reason. When reading a paper, look for the intervals; they tell you what the study has actually demonstrated.

The illusion of precision

Numbers in research papers carry an aura of precision they do not always deserve. A point estimate of 10.27 percent reduction sounds precise; the same finding reported as “somewhere between 4 and 16 percent” feels less so. The latter is closer to honest. Citizens reading research should be wary of point estimates given to two or three decimal places without their accompanying intervals. The decimals are a courtesy of the calculator, not a feature of the world.

The bottom line

Sample size and effect size, taken together with the variability of measurements, determine whether a study has any chance of detecting what it is looking for. Most published studies that fail to replicate were underpowered: too small to reliably detect the effects they reported finding. Statistically significant results from small studies are more likely than they look to be exaggerations or false positives. Confidence intervals tell you both the estimate and its precision; they are usually more useful than p-values alone. As a citizen reader, ask: How many subjects? How big is the claimed effect? Could a study of this size really detect an effect that small with confidence?

What to read or watch next

  • Geoff Cumming, Understanding the New Statistics: Effect Sizes, Confidence Intervals, and Meta-Analysis (Routledge, 2012). A clear book-length treatment of why effect sizes and confidence intervals are more useful than p-values for most purposes.
  • Andrew Gelman and John Carlin, “Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors,” Perspectives on Psychological Science 9 (2014): 641–51. Influential argument for the dangers of underpowered research and the inflation of effect-size estimates.
  • Daniel Läkens, “Calculating and Reporting Effect Sizes,” Frontiers in Psychology 4 (2013): 863. Practical introduction to effect-size statistics; Lakens’s blog (daniellakens.blogspot.com) is also a useful resource.

CHAPTER 6

P-Values, Confidence Intervals, and What They Mean

The p-value is the most ubiquitous and most misunderstood number in modern science. Tens of thousands of research papers each year report p-values; entire careers have been built on producing them; entire reform movements have organized around what they do and do not mean. A citizen reading research will encounter the p-value constantly, and almost every common interpretation of it is wrong. This chapter explains what the p-value actually is, what it is not, and what the working scientific community currently makes of it.

What a p-value actually is

Formally: the p-value is the probability of observing data as extreme as the data you observed, or more extreme, if the null hypothesis (typically, “there is no real effect”) were true. Less formally: imagine that the treatment had no real effect. Imagine running the experiment many, many times under that assumption. In what fraction of those imaginary experiments would you see a result at least as striking as the one you actually got? That fraction is the p-value. A small p-value means: “if there were really no effect, results like the one I observed would be rare.”

The conventional threshold of p < 0.05 means that, if the null hypothesis were true, results at least as striking as those observed would occur less than 5 percent of the time. This threshold was popularized by the statistician Ronald Fisher in the 1920s as a rule of thumb, not a fundamental law. It has since acquired a vastly disproportionate weight in scientific practice. Findings that just clear it get published; findings that just miss it often disappear. In 2016, the American Statistical Association issued a formal statement warning against the over-interpretation of p-values, and in 2019 a coalition of statisticians and editors signed a comment in Nature titled “Scientists rise up against statistical significance,” arguing the threshold has done more harm than good.

What a p-value is not

The p-value is not the probability that the null hypothesis is true. This is the most common misinterpretation, and it is wrong. A p-value of 0.03 does not mean “there is a 3 percent chance the effect isn’t real.” It means “if the effect weren’t real, results this striking would occur 3 percent of the time by chance.” Those are different statements. The probability that the null hypothesis is true depends not only on the p-value but on how plausible the hypothesis was before the study began — the so-called prior probability — which the p-value alone cannot tell you. A study testing an implausible hypothesis with p = 0.04 is not particularly strong evidence; a study testing a well-grounded hypothesis with the same p-value is more compelling.

The p-value is also not a measure of effect size. A p-value of 0.001 does not mean the effect is large or important; it means the result is unlikely under the null hypothesis. With a sufficiently large sample, vanishingly small effects can produce vanishingly small p-values. The p-value tells you about the role of chance, not about the magnitude or importance of what was found. This is why effect sizes and confidence intervals are essential complements.

Finally, the p-value is not a measure of replication probability. A p-value of 0.05 does not mean the result will replicate 95 percent of the time. The actual probability of replication depends on the true effect size, the sample size of the replication, and the proportion of effects in the field that are real to begin with. A field full of barely-significant results from underpowered studies will have a low replication rate even if every individual p-value is computed correctly.

The replication crisis and p-values

Much of the empirical evidence on the replication crisis can be summarized in p-value terms. Across many fields, the distribution of published p-values shows a striking discontinuity at 0.05 — far more papers report results just below the threshold than just above. This pattern is statistically incompatible with the underlying effects being honestly measured; it is a fingerprint of selective reporting, p-hacking, and outright fraud, the topics of Chapter 8. The shape of the curve has been documented by researchers like Uri Simonsohn, Joseph Simmons, and Leif Nelson, who in 2014 published a method called p-curve analysis for diagnosing whether a body of research is reporting honest evidence or selectively reporting whatever crossed the threshold.

Bayesian alternatives

A growing minority of researchers prefer Bayesian statistics, which directly compute the probability of a hypothesis given the data, using both the data and an explicit prior probability. Bayesian methods avoid some of the conceptual problems of p-values but require their own judgment calls (what prior should be assumed?) and have not yet displaced the standard frequentist framework in most applied research. Citizens are unlikely to encounter Bayesian results outside specialized journals, but it is worth knowing that the dispute exists; a single number reporting “p = 0.04” contains less information than scientists themselves would ideally want.

Confidence intervals revisited

Recall from Chapter 5 that a 95 percent confidence interval is a range of values, calculated so that in repeated samples it would contain the true value 95 percent of the time. The intuitive interpretation — “I am 95 percent sure the true value is in the interval” — is, like the intuitive interpretation of p-values, technically wrong but practically not very misleading in most cases. The right way to use confidence intervals is to look at where they fall and how wide they are. An interval entirely above zero means the data support a positive effect, with the bounds telling you how big or small it might plausibly be. An interval that crosses zero means the data are consistent with no effect. A wide interval means the data have not pinned the answer down; a narrow one means they have.

The dance of the means

Geoff Cumming, an Australian statistician, has popularized a useful demonstration he calls the “dance of the means.” If you simulate repeated samples from a population with a known effect, and plot the resulting confidence intervals, you see how much they jump around — even when the underlying truth is fixed. The visualization undoes the intuition that any single study’s estimate is the answer. It is one estimate among many that could have been produced; the next study, run in good faith with the same methods, would produce a slightly different number with a slightly different interval. Reading any single research finding as definitive ignores the fundamental noisiness of the process.

The replacement debate

In recent years, statisticians and methodologists have debated whether to abandon the conventional 0.05 threshold altogether. Some have proposed lowering it to 0.005 for new claims; others have proposed abandoning thresholds and reporting p-values as continuous evidence; others have proposed replacing p-values entirely with Bayesian or confidence-interval-based methods. No consensus has emerged, but a citizen reading research in 2026 should know that the conventional threshold is increasingly contested within the scientific community itself. Findings that depend critically on the difference between p = 0.04 and p = 0.06 should be regarded as preliminary, regardless of which side of the threshold they fall on.

The bottom line

A p-value is the probability of seeing data as extreme as the data observed, assuming the null hypothesis is true. It is not the probability that the null hypothesis is true, or a measure of effect size, or a guarantee of replication. The conventional threshold of 0.05 is a rule of thumb that has acquired more weight than it deserves. Confidence intervals carry more information; both should be reported. A finding’s strength depends on prior plausibility, study size, effect size, and replication, not on the p-value alone. Citizens encountering p < 0.05 should ask: How likely was the hypothesis to begin with? How big is the effect? Has the result replicated?

What to read or watch next

  • Ronald Wasserstein and Nicole Lazar, “The ASA Statement on p-Values,” The American Statistician 70 (2016): 129–33. The American Statistical Association’s formal statement on the limits of p-values; clear, short, and authoritative.
  • Valentin Amrhein, Sander Greenland, and Blake McShane, “Scientists Rise Up Against Statistical Significance,” Nature 567 (2019): 305–7. Comment piece, signed by more than 800 scientists, arguing for retiring the 0.05 threshold.
  • Regina Nuzzo, “Scientific Method: Statistical Errors,” Nature 506 (2014): 150–52. The most widely shared general-audience explanation of why p-values are so misunderstood.
  • Aubrey Clayton, Bernoulli’s Fallacy: Statistical Illogic and the Crisis of Modern Science (Columbia University Press, 2021). A philosophical and historical critique of frequentist statistics, arguing for a Bayesian alternative. Provocative but instructive.

PART THREE

The Replication Crisis and the Response

Why findings sometimes don’t hold up, what practices produce false positives, and how the scientific community is responding

CHAPTER 7

Why Many Findings Don’t Replicate

Replication is the soul of science. A finding that holds up only in the laboratory of the original investigator, by their methods, with their measurements, is not yet a finding the world can rely on. A finding that other researchers can reproduce, in different places with different subjects, is. The discovery of how often modern findings fail to replicate — in psychology, in cancer biology, in social science, in clinical research — is the central methodological story of the past fifteen years. To read research as a citizen in 2026 is to read in the shadow of the replication crisis, and to know what its lessons are.

The reproducibility project

The defining moment came in August 2015, when the Open Science Collaboration — a consortium of 270 researchers coordinated by Brian Nosek and the Center for Open Science — published in the journal Science the results of a four-year project to replicate 100 studies from three top psychology journals. The original studies had reported statistically significant effects in 97 of the 100 cases. The replication studies, conducted with high statistical power and largely the same methods (often with the original authors’ cooperation), reported significant effects in only 36. Effect sizes in the replications were on average about half the size of those in the originals. By multiple criteria — statistical significance, effect-size overlap, subjective judgment of project members — fewer than half of the original findings were successfully reproduced.

The reaction was substantial. Some researchers argued the methodology of the replication project was unfair, that subtle differences between original and replication studies could account for the failures. Others argued that the failures revealed a deep problem with how psychology had been done. Subsequent replication projects in other fields produced similar patterns. The Reproducibility Project: Cancer Biology, attempting to replicate 53 high-impact cancer biology findings, completed 50 replications and found that effect sizes were a fraction of those in the originals. Replication projects in experimental economics found higher rates than in psychology but still substantial failures. The pattern was too consistent to dismiss.

Why does it happen?

There is no single cause of replication failure. Several distinct mechanisms operate, sometimes in combination, to produce findings that do not hold up:

  • Chance. Some statistically significant findings are simply false positives — effects that appeared by chance and would not be found again. With a 5 percent significance threshold, 5 percent of true null hypotheses will produce significant findings by chance. In a literature dominated by underpowered studies, the false-positive rate among published claims is much higher than 5 percent.
  • Effect-size inflation (winner’s curse). As discussed in Chapter 5, underpowered studies that find significant effects systematically overestimate them. Replication studies, with adequate power, find smaller effects — often genuine, but smaller than originally reported.
  • Publication bias. Studies finding significant effects are far more likely to be published than studies finding null results. The published literature therefore overstates the strength of the evidence; replications run with proper power often fail because the published effect was already an overestimate.
  • P-hacking and selective reporting. Researchers who try multiple analyses, multiple measures, or multiple subgroups, and report only what worked, can produce statistically significant findings from data that contain no real effect. Chapter 8 explores these practices in detail. They are common, often unintentional, and the literature does not always disclose them.
  • Real heterogeneity. Some failures of replication occur because the underlying effect is genuine in some contexts but not others, or interacts with subject populations or experimental settings in unexpected ways. A finding that holds among undergraduates may fail to replicate among older adults; an effect documented in one country may not appear in another. Sometimes “replication failure” means “replication in a different context, with different results.”
  • Outright fraud. A small number of replication failures trace to fabricated or manipulated data in the originals. Most replication failures are not fraud, but the prominence of recent fraud cases (Chapter 12) shows that some fraction of the published literature is unreliable for this reason.

How widespread is the problem?

The replication crisis is not uniform across science. The fields hit hardest are those where small samples are typical, effects are subtle, the noise is high, and incentives reward novelty over rigor. Social psychology, education research, and parts of nutrition and biomedical science have all shown serious replication problems. Fields with larger samples, harder measurements, or stronger theoretical constraints — physics, chemistry, much of clinical medicine, parts of economics — have shown fewer problems, though they are not immune. The crisis is most accurately understood as a crisis of certain methodological practices, not of science as such.

Surveys of researchers themselves bear out the picture. A 2016 survey of 1,576 researchers by Nature found that more than 70 percent had failed to replicate at least one other scientist’s results, and more than half had failed to replicate one of their own. Most respondents agreed that there was a significant or slight crisis of reproducibility. The picture from inside the scientific community is not one of business as usual.

Cancer biology and beyond

Some of the most consequential evidence on replication has come from cancer biology. In 2012, Glenn Begley and Lee Ellis published in Nature an account of an attempt by the pharmaceutical company Amgen to replicate 53 “landmark” preclinical cancer studies as a basis for drug development. They were able to confirm only six — a replication rate of about 11 percent. Bayer reported similar problems in its own internal replication efforts. The Reproducibility Project: Cancer Biology, conducted by the Center for Open Science between 2013 and 2021, attempted to replicate 53 high-impact cancer findings and found average effect sizes 85 percent smaller than the originals. The implications for drug development — where unreplicated preclinical findings can lead to expensive failed trials and missed opportunities — are profound.

What the citizen should take from all this

The lesson of the replication crisis is not that science is broken. It is that any single published finding, especially in fields known to have replication problems, deserves more skepticism than the published reputation of “peer-reviewed research” might suggest. A finding that has been independently replicated in multiple settings is much stronger evidence than a single splashy result. A finding from a large, preregistered study with effect sizes adequate to be confident in is much stronger than a small exploratory study reporting a barely-significant effect. The shift in mental model is from “this study found X” to “this study is one piece of evidence about X; what does the broader pattern look like?”

The bottom line

Many published findings, especially in fields with small samples and subtle effects, fail to replicate. The reasons include chance, effect-size inflation in underpowered studies, publication bias, p-hacking, real heterogeneity, and occasional fraud. The crisis is more severe in some fields than others, but it should change how citizens read individual studies. A single published finding is much weaker evidence than the published reputation of peer review suggests; a finding replicated independently is much stronger. Ask not just “did a study find this?” but “how many independent studies have found it, and how consistently?”

What to read or watch next

  • Open Science Collaboration, “Estimating the Reproducibility of Psychological Science,” Science 349 (2015): aac4716. The foundational replication study; technical but worth at least skimming.
  • Stuart Ritchie, Science Fictions (Metropolitan Books, 2020). Chapters on replication, fraud, and bias offer the best general-audience treatment available.
  • Ed Yong, “Psychology’s Replication Crisis Is Real,” The Atlantic (March 14, 2018). Excellent reporting on the state of the field by one of the leading science journalists of the era.
  • C. Glenn Begley and Lee Ellis, “Raise Standards for Preclinical Cancer Research,” Nature 483 (2012): 531–33. The famous Amgen replication attempt and its disturbing findings.

CHAPTER 8

P-Hacking, HARKing, and Other Bad Practices

If many published findings fail to replicate, the natural question is why. Some of the answer is innocent — the statistical issues described in Chapter 7. But part of the answer involves practices that, while sometimes well-intentioned, systematically produce false positive results and inflated effect sizes. These practices have collectively been called “questionable research practices” or QRPs, and the catalog of them is now well documented. None requires conscious dishonesty. A researcher can engage in any of them with the best intentions and still produce publications that will not hold up. Citizens reading research benefit from understanding the practices, both because they help explain why so much evidence is unreliable and because awareness of them changes how to read the methods sections of papers.

P-hacking

P-hacking refers to the family of practices by which a researcher tries multiple analyses on the same data and reports only those that produce statistically significant results. The basic logic is simple: if you test a single hypothesis with a 5 percent significance threshold, the chance of a false positive is 5 percent. If you test 20 hypotheses on the same data, the chance that at least one will produce a false positive is much higher — about 64 percent. If you have flexibility in how to define your variables, what subgroups to examine, what covariates to include, when to stop collecting data, and which outliers to exclude, the effective number of tests can be very large indeed.

In a famous 2011 demonstration, Joseph Simmons, Leif Nelson, and Uri Simonsohn published a paper in Psychological Science titled “False-Positive Psychology,” in which they showed that with relatively modest analytical flexibility — reporting on two dependent variables, adding observations until significance was reached, controlling for gender, and dropping or not dropping a condition — the false-positive rate could rise from a nominal 5 percent to over 60 percent. The same paper humorously demonstrated, using these techniques, that listening to the Beatles song “When I’m Sixty-Four” made participants younger — a result that, taken seriously, would imply that listening to music alters chronological age. The paper was a turning point in awareness.

Common forms of p-hacking

  • Optional stopping. Continuing to collect data until the analysis reaches significance, then stopping. This inflates false-positive rates because results that would have failed with the original sample size can succeed by adding subjects who happen to push the result over the threshold.
  • Outcome flexibility. Measuring multiple outcomes and reporting only those that produced significant results. If a study measured aggression in five different ways, and only one showed a significant effect, reporting only the one that worked exaggerates the strength of the evidence.
  • Subgroup fishing. Slicing the data into subgroups (gender, age, region, etc.) and reporting only the slices that produced significant effects, often with a post-hoc story about why that group was the interesting one.
  • Covariate manipulation. Trying many combinations of statistical controls and reporting the analysis that produces the cleanest result, often without disclosing the alternatives that were tried.
  • Outlier handling. Excluding observations as “outliers” when their inclusion weakens the result, including them when their exclusion would weaken the result, with the choice driven by what makes the headline finding work.

HARKing

HARKing — Hypothesizing After the Results are Known — is the practice of running an exploratory analysis, finding an interesting pattern, and then writing the paper as if you had predicted that pattern in advance. This is a particularly insidious problem because nothing in the published paper looks wrong. The pattern is real, in the data; the analysis is sound; the conclusion follows. What is missing is the disclosure that the hypothesis was generated from the same data that supposedly tested it. From the reader’s perspective, the paper looks like a successful confirmatory study. From a scientific perspective, it is exploratory — and exploratory findings are far more likely to be false positives than confirmatory ones. Without preregistration (Chapter 9), the difference between confirmatory and exploratory analyses cannot be reliably detected from the published paper.

The garden of forking paths

Andrew Gelman and Eric Loken, in a 2013 paper called “The Garden of Forking Paths,” pointed out that p-hacking does not require conscious dishonesty. A single researcher, working in good faith on a single dataset, faces dozens of small analytical choices: how to code variables, which covariates to include, how to handle missing data, where to set thresholds, which subgroups to examine. Each choice is reasonable; many are not specified in advance; the researcher makes them naturally as the analysis unfolds. But the choices are not made in a vacuum. They are made in light of what the data are showing, with intuition guided by which choices produce cleaner-looking results. The cumulative effect, even without any deliberate fishing, is to produce a path through the analysis that selects favorable findings. The practice is not fraud — it is the way reasonable people analyze data when no one has forced them to commit in advance.

Publication bias and the file drawer

Independent of any individual researcher’s practices, the broader incentive structure of academic publication produces systematic distortion. Studies finding significant results are much more likely to be published than studies finding null results. The studies that find nothing tend to languish in researchers’ file drawers — hence the term file-drawer problem. The published literature therefore overstates the strength of evidence on most questions, sometimes by enormous margins. Meta-analyses that try to combine published evidence inherit this bias unless they explicitly correct for it. The phenomenon was identified as far back as 1959 (by the statistician Theodore Sterling) and has been confirmed repeatedly since.

How can a citizen tell?

It is hard to detect p-hacking or HARKing from a single paper without forensic statistical analysis. There are some signals: a paper that reports many marginal findings (p-values clustered just below 0.05), that emphasizes subgroup results, or that makes its hypothesis sound suspiciously well-fitted to the data. But the most reliable signal is whether the study was preregistered — the topic of the next chapter — and whether replications have followed. Citizens reading a paper that was not preregistered and has not been replicated should treat the headline finding as preliminary, no matter how prestigious the journal.

The bottom line

Researchers do not need to commit fraud to produce findings that will not replicate. The combination of analytical flexibility (p-hacking), retrospective hypothesis-fitting (HARKing), and publication bias is enough. None of these practices necessarily reflects bad faith; they emerge naturally from the way research is conducted and rewarded. Citizens reading a single paper cannot easily detect them. The best protection is to give weight to findings that have been preregistered or independently replicated, and to discount the rest accordingly.

What to read or watch next

  • Joseph Simmons, Leif Nelson, and Uri Simonsohn, “False-Positive Psychology,” Psychological Science 22 (2011): 1359–66. The original demonstration of how analytical flexibility produces false positives. Short and clear; written for a non-statistical audience.
  • Andrew Gelman and Eric Loken, “The Garden of Forking Paths” (2013). Working paper available at gelman.com; explains why honest researchers can produce non-replicable findings without fraud.
  • Norbert Kerr, “HARKing: Hypothesizing After the Results are Known,” Personality and Social Psychology Review 2 (1998): 196–217. The original paper coining the term and describing the problem.
  • Data Colada (datacolada.org). The blog of Simmons, Nelson, and Simonsohn; ongoing analysis of methodological problems and outright fraud, often in plain language. The four-part 2023 series on Francesca Gino is essential reading on what fraud detection actually looks like.

CHAPTER 9

Preregistration, Open Data, and the Reform Movement

The replication crisis prompted serious reform. Over the past decade, a movement — sometimes called the Open Science movement, the Credibility Revolution, or simply Reform — has pushed the academic community to adopt practices that make research more reliable, more transparent, and more verifiable. The reforms have not solved the problems they were designed to address, but they have made measurable progress, and they have shifted the methodological norms of several major fields. Citizens reading research in 2026 will increasingly encounter the marks of these reforms: preregistration tags, open data badges, registered reports, replication studies. Understanding what they mean is part of reading research well.

Preregistration

The simplest and most important reform is preregistration: the practice of publicly committing to a study’s hypotheses, design, and analysis plan before the data are collected (or, for analyses of existing data, before the data are examined). The commitment is recorded with a time-stamp on a registry such as the Open Science Framework (osf.io), AsPredicted (aspredicted.org), or ClinicalTrials.gov. When the study is later published, the actual analyses can be compared to the preregistered plan. Deviations are not forbidden — sometimes they are necessary — but they have to be disclosed.

Preregistration addresses many of the practices described in Chapter 8 in a single move. P-hacking becomes much harder when the analysis plan is fixed in advance. HARKing becomes detectable when the registered hypothesis differs from the published one. Publication bias is reduced because preregistered null results have a clearer claim to publication. Preregistration does not eliminate questionable practices entirely — researchers can register vague plans, deviate without disclosure, or simply abandon registered studies whose results disappoint — but it raises the bar substantially.

Registered reports

A more ambitious reform is the registered report: a publication format in which researchers submit their study design and analysis plan for peer review before collecting data. If the design is approved, the journal commits in advance to publishing the results, regardless of whether they confirm the hypothesis. Registered reports formally separate the question of whether a study is well-designed from the question of whether it produced an interesting result — and they make publication independent of the result. Approximately 300 journals now offer registered reports, including some of the highest-profile journals in psychology, neuroscience, and a growing range of fields.

Studies of registered reports have found that they produce a much higher proportion of null findings than the regular literature — around 40 to 50 percent, compared with under 5 percent in conventional papers — and that effects, when found, are typically smaller. This pattern is what one would expect if conventional publication strongly selects for positive findings. Registered reports do not solve every problem, but they offer the cleanest current solution to the file-drawer problem and the perverse incentives around result-driven publication.

Open data and code

A complementary reform is the open sharing of data, materials, and analysis code. Many journals now require, or strongly encourage, that authors deposit their data in public repositories so that other researchers can verify, replicate, or extend the analyses. Open code allows readers to see exactly what analyses were performed and to detect deviations from preregistered plans. The Center for Open Science’s Open Science Framework (osf.io) hosts much of this material. Several high-profile fraud cases of the past decade — most notably the work of Francesca Gino, exposed in 2023 by the Data Colada team — have been resolved precisely because data sleuths could examine raw data files and detect inconsistencies that would have been invisible from the published papers alone.

Replication initiatives

Beyond individual reforms, a number of organized replication initiatives now exist. Many Labs is a series of large-scale collaborations in which multiple research groups attempt the same studies, often producing more definitive evidence than any single laboratory could. Replication Markets is a project in which researchers bet on which findings will replicate, with the resulting market prices providing a quick guide to which published claims researchers themselves regard as likely to hold up. The Reproducibility Projects in psychology and cancer biology, mentioned in Chapter 7, are part of this larger pattern. The cumulative effect is a slowly emerging map of which fields and which kinds of findings are reliable and which are not.

The Center for Open Science

Much of the institutional infrastructure of the reform movement has been built by a single nonprofit, the Center for Open Science, founded in 2013 by University of Virginia psychologist Brian Nosek. The Center hosts the Open Science Framework, manages preregistration tools, coordinates large replication projects, and develops publishing standards (the Transparency and Openness Promotion guidelines) that hundreds of journals have adopted. The Center for Open Science is a remarkable example of a small organization with a clear mission accelerating the reform of an entire field; citizens curious about how academic culture is changing should know that it exists.

How the field is changing

The reforms are not universally embraced. Some senior researchers regard them as bureaucratic intrusions; some early-career researchers worry that strict preregistration disadvantages those whose findings turn out to be modest. The pace of reform varies dramatically across fields: psychology, neuroscience, and parts of biomedical science have moved faster than economics, sociology, or many medical specialties. Top journals have adopted reforms more quickly than middle-tier ones. Sample sizes are larger than they were a decade ago, p-curve analyses are reported, registered reports are growing, but a great deal of the published literature still operates under the older norms.

What this means for reading

A practical implication for citizen readers is that the date and the source of a study now matter even more than they used to. A study published in 2024 in a journal that has adopted strict reporting standards, with preregistration and open data, is on average more reliable than a study published in 2010 in the same journal under older norms. A study from a field with a strong reform tradition (recent psychology, much of clinical medicine) is more reliable than one from a field where reform has lagged. None of this is a guarantee, but the trends are real and worth noting when weighing evidence.

The bottom line

The replication crisis has prompted a serious reform movement. Preregistration commits researchers to a plan before they see results; registered reports decouple publication from outcomes; open data and code allow independent verification; replication initiatives produce systematic evidence on which findings hold up. The reforms are imperfect, uneven across fields, and slowly adopted, but they have measurably improved the reliability of new research. Citizens reading a study should ask: was it preregistered? Are the data and code available? Has the finding been independently replicated? Affirmative answers should raise confidence; negative answers should lower it.

What to read or watch next

  • Brian Nosek et al., “Promoting an Open Research Culture,” Science 348 (2015): 1422–25. The Transparency and Openness Promotion (TOP) guidelines that hundreds of journals have since adopted.
  • Center for Open Science (cos.io). Hub of the reform movement; the Open Science Framework (osf.io) is where most preregistration and data-sharing happens.
  • Christopher Chambers, The Seven Deadly Sins of Psychology: A Manifesto for Reforming the Culture of Scientific Practice (Princeton University Press, 2017). Insider account of the problems and the reform movement, written by one of the leading proponents of registered reports.
  • Marcus Munafò et al., “A Manifesto for Reproducible Science,” Nature Human Behaviour 1 (2017): 0021. Comprehensive overview of the reform agenda by an interdisciplinary group of researchers.

PART FOUR

Kinds of Research and How to Read Them

Medical research and clinical trials, social science, and the special features of economics, education, and policy research

CHAPTER 10

Medical Research and Clinical Trials

Medical research is among the highest-stakes work the scientific community produces. The conclusions inform what drugs you take, what surgeries you receive, what screening tests you undergo, and what advice your doctor gives you. The standards for medical research are correspondingly high — in some respects higher than in any other field. The methodologies are well-developed, the regulatory oversight is substantial, and the financial costs of getting things wrong are enormous. None of this prevents medical research from making mistakes, sometimes serious ones. Reading medical research as a citizen requires understanding both its formal strengths and its real-world failures.

The structure of clinical trials

Clinical trials of new drugs and treatments are typically conducted in phases, each with a different purpose:

  • Phase I. Small studies, typically 20–80 healthy volunteers or sometimes patients, designed to assess safety and find the appropriate dose. Phase I results tell you almost nothing about whether the drug works.
  • Phase II. Mid-sized studies, typically 100–300 patients, designed to assess preliminary efficacy and further evaluate safety. Phase II results suggest whether a drug is worth pursuing but cannot reliably establish that it works.
  • Phase III. Large randomized controlled trials, typically thousands of patients, designed to definitively establish efficacy and safety. Phase III results are the basis on which drugs are approved.
  • Phase IV. Post-approval surveillance, sometimes called post-marketing studies, designed to detect rare side effects or long-term outcomes that smaller trials cannot capture.

A citizen encountering a news story about a “promising new treatment” should always ask which phase the supporting evidence comes from. A Phase I or II result, however striking, is preliminary; the majority of drugs that show promise in early trials fail to demonstrate efficacy in Phase III. The history of cancer research is full of treatments that worked beautifully in mice or in small early-stage human trials and disappointed in large randomized trials. The phase distinction is not bureaucratic. It is the difference between a hypothesis worth investigating and a finding worth acting on.

Endpoints, surrogates, and what was measured

A clinical trial measures something. Sometimes it measures the outcome you actually care about — do patients live longer, have fewer heart attacks, recover function, feel better. Sometimes it measures a surrogate — a biomarker, lab value, or imaging finding presumed to track the real outcome but that may or may not. Surrogate endpoints are tempting because they can be measured faster and in smaller trials, but they have a long history of misleading. Drugs that improve cholesterol numbers do not always reduce heart attacks; drugs that shrink tumors do not always extend life. The 1980s saw multiple antiarrhythmic drugs approved on the basis of their effect on a surrogate (suppression of certain abnormal heartbeats) only to discover, when long-term mortality was studied, that they actually increased deaths. Reading a clinical trial means asking what was measured and whether it was the thing that mattered.

Absolute versus relative risk

Few statistical confusions in medical research are as consequential as the difference between absolute and relative risk. Suppose a drug reduces the risk of heart attack from 2 percent to 1 percent over five years. The relative risk reduction is 50 percent — the drug halves your risk. The absolute risk reduction is 1 percentage point — you went from 98 percent chance of no heart attack to 99 percent chance. Both numbers are correct. Both describe the same finding. But they communicate very different impressions of the magnitude of the benefit, and clinical trials and press releases routinely report only the relative figure because it sounds more dramatic. Citizens reading about “a 50 percent reduction in risk” should ask: 50 percent of what? A 50 percent reduction in a 1 percent risk is a half-percentage-point shift. Acting on it may still be reasonable, but the magnitude is not what the headline implies.

A related figure, the number needed to treat (NNT), conveys the same information more usefully. The NNT is the number of patients who must receive the treatment for one to benefit. In the example above, you would need to treat 100 patients for five years to prevent one heart attack — NNT = 100. NNTs are sometimes published in clinical trial reports and almost never in news stories. They are the most informative single number for thinking about whether a treatment is worth taking, and citizens learning to ask for them gain a real edge in interpreting medical research.

Industry funding

Most large clinical trials are funded by pharmaceutical or device companies that stand to profit from a positive result. This is not in itself disqualifying — industry funds the majority of drug development, and most regulatory frameworks are designed to manage the conflict — but it is a relevant fact. Studies have repeatedly shown that industry-funded trials are more likely to report results favorable to the sponsor’s product than independent trials of the same questions. The mechanisms include study design choices (comparing the new drug to placebo rather than to the best existing treatment, or to a low-dose comparator), publication bias (unfavorable results not being published), and selective reporting of outcomes. The Cochrane Collaboration’s reviews routinely identify these patterns and try to correct for them.

FDA approval and what it means

In the United States, drugs and devices reach the market through approval by the Food and Drug Administration. FDA approval is a serious bar, requiring substantial evidence of safety and efficacy from well-designed trials. It is also not a guarantee that a drug works, or that the version that reaches you will. Approvals can be based on surrogate endpoints; they can rely on a single positive trial after multiple negative ones; accelerated-approval pathways allow drugs to reach the market on preliminary evidence with the requirement that confirmatory trials follow. The FDA’s standards are higher than those of regulators in many other countries, but they are not infallible. The history of drugs withdrawn from the market after approval — Vioxx, fenfluramine-phentermine, and others — reminds us that approval is the start, not the end, of the evidence-gathering process.

The Cochrane Collaboration

For most clinical questions, the most reliable summary of the evidence is a Cochrane systematic review. Cochrane is an international nonprofit, founded in 1993 and named for the British epidemiologist Archie Cochrane, that produces methodologically rigorous reviews of medical evidence. Cochrane reviews follow standardized methods, register their protocols in advance, and update as new evidence accumulates. They include plain-language summaries written for non-specialists. A citizen wondering whether a treatment works — antibiotics for bronchitis, vitamin D for prevention, fish oil for cardiovascular disease — will usually get a more honest answer from a Cochrane review than from any single study or any news article.

Reading a clinical trial: a checklist

When reading a clinical trial, citizens can apply a small checklist:

  • Was it randomized? Look for explicit description of randomization.
  • Was it blinded? Single-blind, double-blind, or open-label?
  • What was the sample size, and were enrollment numbers maintained throughout?
  • What was the primary endpoint, and was it the outcome that matters or a surrogate?
  • Were absolute risk numbers reported, or only relative?
  • Who funded the study, and were authors’ conflicts disclosed?
  • Was the trial preregistered, and were the published analyses consistent with the registration?
  • Has the finding been replicated in independent trials, and is it consistent with prior evidence?

The bottom line

Medical research has the most developed methodology of any field, with phased clinical trials, FDA oversight, and the Cochrane infrastructure for systematic review. It is also subject to industry funding bias, surrogate-endpoint pitfalls, and the routine confusion of relative and absolute risk. Citizens reading medical research should attend to phase, endpoint, absolute risk reduction, funding source, and replication. For most clinical questions, a Cochrane systematic review is more reliable than any single study or any news article.

What to read or watch next

  • Cochrane Collaboration (cochrane.org). Plain-language summaries of systematic reviews on most clinical questions; the first place to look for evidence on whether a medical treatment works.
  • Ben Goldacre, Bad Pharma: How Drug Companies Mislead Doctors and Harm Patients (Faber & Faber, 2012). Detailed account of the systematic problems in pharmaceutical research and what reforms are needed; readable and angry but well-sourced.
  • Marcia Angell, The Truth About the Drug Companies (Random House, 2004). Former editor-in-chief of the New England Journal of Medicine on what is wrong with pharmaceutical research and marketing.
  • Trisha Greenhalgh, How to Read a Paper (Wiley-Blackwell, 6th ed., 2019). Especially the chapters on clinical trials and on systematic reviews.
  • Vinay Prasad, Malignant: How Bad Policy and Bad Evidence Harm People with Cancer (Johns Hopkins University Press, 2020). An oncologist’s critique of the use of weak evidence in cancer drug approval; technical but accessible.

CHAPTER 11

Social Science Research

Social science research — the scholarly study of human behavior, social organization, attitudes, and institutions — covers an enormous range of disciplines and methods. Psychology, sociology, political science, anthropology, social epidemiology, parts of economics and education research all fall under the umbrella. The substantive topics are often the ones that matter most to citizens: how children develop, how families function, how communities work, what makes people happy, how prejudice operates, what drives crime, what shapes political opinions. Unfortunately, the methodological challenges in social science are severe, and the past decade has been particularly hard on its credibility. Reading social science research as a citizen requires both interest in the questions and skepticism about the answers.

Why social science is harder than it looks

Social phenomena are inherently difficult to study. Human behavior is variable, context-dependent, and shaped by countless factors that researchers cannot control. The constructs of interest — “intelligence,” “aggression,” “well-being,” “prejudice” — are abstract and often contested, requiring measurements that translate them into something countable. Randomized experiments, where they are possible at all, often involve artificial situations whose generalizability to real life is uncertain. Observational studies face all the challenges of confounding discussed in Chapter 4. Effect sizes, even when real, are often small, requiring large samples to detect reliably.

On top of these inherent difficulties, social science as a field has historically had cultural norms that worsened the problem: small sample sizes were tolerated, replication was undervalued, exploratory and confirmatory analyses were not clearly distinguished, novelty was rewarded over rigor. The replication crisis hit social psychology particularly hard precisely because these norms had become so entrenched. The good news is that the reform movement has been most active in social psychology and adjacent fields, and recent work is, on average, considerably more rigorous than work from a decade or two ago.

Famous failed replications

A short tour of social-science findings that turned out to be unreliable will help citizens calibrate. The “power posing” research, claiming that adopting confident postures for two minutes raised testosterone and lowered cortisol, was widely popularized through a TED talk and a book; subsequent replications failed to find the hormonal effects, and one of the original co-authors, Dana Carney, publicly disowned the work. “Ego depletion,” the theory that willpower is a finite resource that runs down with use, was supported by hundreds of studies but failed in large preregistered replications. “Stereotype threat,” the claim that reminding minorities of stereotypes about their group degrades their test performance, has had a mixed replication record — some effects survive, others do not, and the magnitude is much smaller than originally claimed. “Facial feedback,” the claim that holding a pen in your teeth makes you find cartoons funnier (because the posture activates smile muscles), failed to replicate in a multi-laboratory project. The “Marshmallow test,” famous for predicting life outcomes from a four-year-old’s ability to delay gratification, was largely explained away as a proxy for socioeconomic background when the original sample was extended.

Not all of these findings were entirely wrong, and most of the original researchers were honest, hardworking professionals doing their best. But each finding had circulated in popular culture as if it were settled science, and each turned out to be much weaker than its public reputation suggested. Citizens drawing on social-science findings should remember that famous results are not automatically reliable results.

Survey research and self-report

Much of social science depends on what people say about themselves: their attitudes, behaviors, experiences, beliefs. The methodology of survey research is well-developed (and shares many issues discussed in the polling guide), but self-report has well-known limits. People misremember; they tell interviewers what they think the interviewer wants to hear; they answer questions whose meaning they have not actually understood; they describe themselves in ways that are not how they actually behave. Studies that match self-reported behavior to objective measures — hours slept, money spent, time exercised — typically find substantial and systematic discrepancies. None of this makes self-report useless, but it should temper confidence in findings that depend entirely on what subjects said about themselves.

WEIRD samples

In 2010, Joseph Henrich, Steven Heine, and Ara Norenzayan published a paper called “The Weirdest People in the World?” pointing out that an enormous fraction of psychological research uses samples from Western, Educated, Industrialized, Rich, and Democratic societies — hence WEIRD — and within those, often samples of undergraduates at a single university. The empirical question they raised is whether findings from these samples generalize to other populations. The answer, in many domains they reviewed, was: not as much as one might hope. Visual perception, moral reasoning, fairness norms, and other supposedly universal human traits showed considerable cross-cultural variation. The implication for citizens is to ask, when reading a social-science finding: who exactly was studied, and to whom is the finding being generalized?

Effect sizes in social science

Even when social-science findings replicate, the effect sizes are often modest. Most interventions in education, parenting, public health, and policy produce small changes in average outcomes. This does not mean the interventions are worthless — small effects, applied across millions of people, can matter — but it does mean that headlines describing “powerful” effects are usually overstating. A useful habit when reading a social-science finding is to convert the reported effect size into a concrete prediction. If a study reports that an intervention improves outcomes by half a standard deviation, what fraction of subjects who would have failed now succeed? If a study reports a 10 percent increase, 10 percent of what — of an absolute rate, of an outcome that is rare to begin with, of a measure most readers cannot interpret? Concretizing the magnitude usually deflates the interpretation.

The role of theory

Social science is theoretically pluralistic to a degree most natural sciences are not. Psychologists, sociologists, economists, and anthropologists studying the same phenomena often start from different theoretical frameworks and arrive at different conclusions. This is not a flaw, but it means that any single finding has to be read in the context of which framework it assumes and what the alternative readings might be. A finding that depression is partly genetic does not exclude social or environmental explanations; a finding that crime tracks economic conditions does not exclude cultural ones. A citizen confronted with a social-science finding should ask not only “is this true?” but “what else is also true that this finding does not capture?”

The bottom line

Social science addresses the questions that often matter most to citizens but works under particularly difficult methodological conditions. Many widely publicized findings have failed to replicate. Self-report has limits; samples are often unrepresentative; effect sizes are typically smaller than headlines suggest; theoretical frameworks shape interpretations. Reading social science well means engaging with both the substance and the methodology, asking who was studied, what was measured, and how robustly the finding has held up. The reform movement has improved practice, but skepticism about any single splashy finding remains warranted.

What to read or watch next

  • Joseph Henrich, Steven Heine, and Ara Norenzayan, “The Weirdest People in the World?” Behavioral and Brain Sciences 33 (2010): 61–135. Influential paper on the WEIRD sample problem; technical but the executive summary is accessible.
  • Stuart Ritchie, Science Fictions (Metropolitan Books, 2020). Particularly thorough on the social-science replication problems; case studies of power posing, ego depletion, and others.
  • Lee Jussim, Social Perception and Social Reality: Why Accuracy Dominates Bias and Self-Fulfilling Prophecy (Oxford University Press, 2012). A counterweight to the popular impression that social cognition is dominated by bias; rigorously argued.
  • Society for the Improvement of Psychological Science (improvingpsych.org). Professional organization devoted to methodological reform in psychology; the conferences and resources are good entry points for citizens curious about the state of the field.

CHAPTER 12

Economics, Education, and Policy Research

A great deal of the research that affects public policy comes from a small cluster of fields: economics, education, public health, and the various interdisciplinary policy specialties. These fields share certain features that distinguish them from natural science and from psychology. They are usually concerned with questions that cannot be addressed by laboratory experiments. They use sophisticated statistical methods to extract causal inferences from observational data. They produce findings that often inform government policy, sometimes affecting millions of people. And they are conducted in a heavily politicized environment, where partisans on every side are eager to use research to support their preferred positions. Reading research in these fields requires understanding both the technical methods and the political context.

The credibility revolution

Economics, in particular, has gone through what its practitioners call the credibility revolution. Beginning in the 1990s, economists — led by figures like Joshua Angrist, David Card, Guido Imbens, and Esther Duflo — increasingly emphasized research designs that could actually identify causal effects. They championed natural experiments (as discussed in Chapter 4), instrumental-variable techniques, regression discontinuity designs, difference-in-differences analyses, and randomized field experiments in development economics. The shift dramatically improved the reliability of economic causal claims. By 2019, the Nobel Memorial Prize in Economic Sciences had been awarded to multiple architects of this approach, recognizing both the technical contributions and the cultural change. Economics today is methodologically more credible on causal questions than it was a generation ago, though it has its own ongoing problems.

The minimum wage debate

A useful illustration of how policy research evolves is the long debate about the minimum wage. The textbook prediction was that raising the minimum wage above market levels would reduce employment, particularly among low-wage workers. In 1994, David Card and Alan Krueger published a study comparing fast-food restaurant employment in New Jersey, which raised its minimum wage, to neighboring Pennsylvania, which did not. They found no employment loss, contradicting the textbook prediction. The methodology — a difference-in-differences natural experiment — was novel, and the conclusion controversial. Subsequent studies using a variety of methods have produced varying results: some find small employment effects, others find none, a few find substantial effects. Meta-analyses synthesizing the evidence have generally found small or near-zero effects from moderate minimum-wage increases, though the picture changes for very large increases. The debate is not settled, but the substantive conversation has moved well beyond the textbook prediction.

The episode illustrates several lessons. Economic theory provides predictions, but the data sometimes refuse to cooperate, and well-designed empirical work can overturn what looked like settled doctrine. A single high-quality study (Card-Krueger) can shift a debate, but the eventual answer requires many studies using many methods. Conclusions about “the” effect of the minimum wage tend to be wrong because effects vary across contexts, magnitudes, and time horizons. And finally, results have political consequences — economists who arrive at inconvenient conclusions for one political faction tend to be praised by the other, and vice versa, regardless of methodological merit. Reading policy research with awareness of the politics is part of reading it well.

Education research

Education research has been particularly hard for citizens to interpret because the stakes are high, the methods are often weak, and the political pressures are enormous. Different fads sweep the field every few years — whole-language reading instruction, growth mindset interventions, social-emotional learning curricula, no-zero grading, value-added teacher evaluation — each backed by some studies and contested by others. Real, robust effects in education research are rare. The largest meta-analyses of educational interventions, such as those compiled by John Hattie, find that almost everything has positive effect sizes (because most interventions are tested by comparing them to no intervention at all), making it hard to distinguish what really works from what merely sounds plausible.

A useful single rule for educational research: be skeptical of dramatic findings from small studies, and look for replication. Carol Dweck’s growth mindset research generated enormous policy enthusiasm before large preregistered replications by independent teams found much smaller effects than the original work, with effects concentrated in specific contexts and student populations. Project-based learning, character education, mindfulness curricula, and many similar interventions have shown the same pattern: enthusiastic initial findings, modest follow-up evidence, ambiguous policy implications. Many education-research reformers are now pushing for the kind of rigorous methods that have transformed development economics, but the field is moving slowly.

Public health and policy

Public health research operates at the intersection of medicine and social science, sharing the strengths and weaknesses of both. Some public health questions can be addressed with randomized trials — of vaccines, of screening tests, of behavioral interventions. Many cannot, and rely on observational data. Smoking-and-lung-cancer evidence accumulated over decades through case-control studies, cohort studies, and biological work, eventually producing a consensus strong enough to guide policy. Other public health questions — the effects of school closures during the COVID-19 pandemic, the optimal alcohol consumption level, the role of dietary salt — remain contested despite extensive research. Citizens trying to read public health research benefit from looking at the convergence (or divergence) across multiple study designs and from organizations like the CDC, the WHO, and Cochrane that try to synthesize the evidence.

Think tanks and advocacy research

Much of what citizens encounter as “research” in the policy space does not come from academic institutions at all. It comes from think tanks — some scrupulously nonpartisan, some explicitly aligned with political movements — and from advocacy organizations that produce reports designed to support their causes. A think-tank study can be excellent, drawing on serious researchers and rigorous methods. It can also be a position paper with citations, designed to give the appearance of evidence to a foregone conclusion. Citizens reading such reports should ask the same questions they would of any research — what was the design, what was measured, how big was the effect, what would falsify it — and should additionally consider what the institutional perspective of the source predicts about the conclusions.

How to read policy research

A few practical habits help with policy research specifically:

  • Look for the design behind the headline number. If the claim is causal, what kind of evidence supports the causal claim?
  • Ask whether the effect, if real, is large enough to matter at the scale of the proposed policy.
  • Consider whether the studied context resembles the context to which the policy would be applied.
  • Look for replication — a single splashy finding from one research team is much weaker than convergent findings from multiple teams using different methods.
  • Identify the funder and the institutional perspective, without assuming bias — but use the information.
  • Where possible, prefer reviews and meta-analyses to single studies; they reflect the broader pattern of evidence rather than the noise of any individual paper.

The bottom line

Economics has been transformed by the credibility revolution, with natural experiments and quasi-experimental designs producing more reliable causal claims than older approaches. Education research remains harder to interpret, dominated by enthusiasm for interventions whose effects often shrink under rigorous study. Public health spans both extremes. Think tanks and advocacy organizations produce work of varying quality; citizens should ask the same methodological questions of these reports as of academic research. Across the policy domain, look at convergence across studies, magnitude of effects, contextual fit, and the broader institutional pattern of who funds what and why.

What to read or watch next

  • Joshua Angrist and Jörn-Steffen Pischke, “The Credibility Revolution in Empirical Economics,” Journal of Economic Perspectives 24 (2010): 3–30. Accessible overview of how economic methods improved over the prior two decades.
  • Esther Duflo and Abhijit Banerjee, Poor Economics: A Radical Rethinking of the Way to Fight Global Poverty (PublicAffairs, 2011). Application of randomized experiments to development questions; both substantive and methodological.
  • John Hattie, Visible Learning (Routledge, 2009). Synthesis of educational meta-analyses; useful as a reference even if some of its specific effect-size estimates have been contested.
  • Dani Rodrik, Economics Rules: The Rights and Wrongs of the Dismal Science (W. W. Norton, 2015). What economic research can and cannot tell us about policy questions; balanced and accessible.

PART FIVE

Where Science Is Settled and Where It Isn’t

How scientific consensus forms, examples of settled and contested questions, and how to read science journalism

CHAPTER 13

How Scientific Consensus Forms

Citizens often hear, in public debates, the phrase “the science says” or “there is a scientific consensus on X.” Sometimes the phrase is used to invoke real and well-established findings; sometimes it is used to brush aside legitimate uncertainty; sometimes it is used to manufacture a sense of authority where none exists. Knowing what scientific consensus actually is, how it forms, and what it is and is not evidence of, is essential to reading science discourse intelligently. This chapter takes the question seriously.

What consensus means

Scientific consensus is the working agreement of most researchers in a field about what the evidence currently supports. It is a sociological fact about the community of researchers, not a logical proof. Consensus emerges when many independent investigators, using a variety of methods, arrive at compatible conclusions over time. It strengthens when the underlying mechanisms are understood, when the empirical pattern survives different ways of measuring and modeling, when researchers from different theoretical orientations agree, and when contrary evidence is investigated and explained or accommodated. Consensus weakens when the evidence rests on a single methodology or research group, when contrary evidence accumulates without being addressed, or when the field’s incentives discourage challenges to received views.

Consensus is not unanimity. There are essentially no important scientific questions on which every credentialed researcher agrees. There are still scientifically credentialed people who reject the link between HIV and AIDS, who deny human-caused climate change, who doubt the safety of vaccines, who reject evolution. Their existence does not mean the questions are open. It means science, like every human enterprise, contains outliers. The right question is not “does everyone agree?” but “what does the relevant body of evidence support, as judged by the people who have devoted careers to studying it?”

How consensus is established

Consensus is established through processes that operate over years or decades:

  • Convergent evidence. Multiple independent lines of evidence pointing to the same conclusion. Climate change is supported by temperature records, ocean heat content, sea level, ice cover, ecological shifts, atmospheric CO2 concentrations, and the basic physics of greenhouse gases — all converging.
  • Mechanism. A theoretical understanding of why the empirical pattern exists. Smoking causes lung cancer is supported by epidemiology and by laboratory work showing how tobacco compounds damage cells.
  • Replication. Findings reproduced by independent teams in different settings using different methods. As discussed in earlier chapters, replication is the workhorse of consensus formation.
  • Engagement with criticism. A field whose practitioners take contrary evidence seriously, address it, and update their conclusions accordingly is producing trustworthy consensus. A field that ignores or punishes critics is producing something else.
  • Synthesis through formal review. Bodies like the National Academies, the Intergovernmental Panel on Climate Change, the Cochrane Collaboration, and (in narrower domains) FDA and CDC advisory panels formally synthesize the evidence and produce reports that represent the considered judgment of the field.

How consensus changes

Consensus does change, and this is one of the great strengths of science. The history of science is a history of conclusions once held confidently being revised in the light of new evidence. The medical consensus on the cause of peptic ulcers shifted in the 1980s and 1990s after Barry Marshall and Robin Warren demonstrated that most ulcers are caused by Helicobacter pylori bacteria, contradicting the previous view that ulcers were caused primarily by stress and stomach acid. The geological consensus shifted in the 1960s when plate tectonics replaced the older fixed-continent model. The astronomical consensus shifted in the 1920s when galaxies beyond the Milky Way were demonstrated. The dietary consensus has shifted multiple times on the role of dietary fat in heart disease. The clinical consensus on hormone replacement therapy shifted in the early 2000s when the Women’s Health Initiative trial found that the therapy increased certain risks, contradicting earlier observational evidence.

In each case, the change took longer than ideally it should have. Senior researchers who had built careers on the older view sometimes resisted; institutional processes were slow; new evidence had to accumulate to overwhelming weight before the field shifted. This is not a flaw of science; it is a feature. A field that changed its consensus on every preliminary finding would be useless. A field that resisted change indefinitely would be ossified. The slow but eventual updating, driven by accumulating evidence, is what makes scientific consensus worthy of citizen trust.

When consensus deserves deference

Consensus that has formed through the processes described above — convergent evidence from multiple methods, mechanistic understanding, replication, engagement with criticism, formal synthesis — deserves substantial weight from citizens. The alternative is to imagine that you, with limited time and no specialized training, can reach better-informed conclusions than the working community of researchers who have devoted careers to the question. This is occasionally true — sometimes outsiders see what insiders miss — but it is uncommon. A citizen who routinely concludes that the consensus is wrong on questions where they have no special expertise is probably fooling themselves about something.

The legitimate citizen response to consensus is not to defer mindlessly but to weight it. A high-confidence consensus on a question that has been intensely studied (the safety of measles vaccines, the existence of human-caused climate change, the age of the earth, the link between smoking and lung cancer) deserves very high weight. A more tentative consensus on a question that is harder to study (the optimal blood pressure target for elderly patients, the cognitive effects of moderate alcohol consumption, the long-term economic effects of trade liberalization) deserves significant but more provisional weight. A weak emerging consensus on a contested question deserves attention but not commitment. Treating all consensuses as if they had the same weight is itself a mistake.

Manufactured controversy

In some prominent public disputes, the appearance of scientific controversy has been deliberately manufactured by interested parties. The tobacco industry pioneered the technique in the 1950s and 1960s, funding research designed to muddy the public picture even as the internal scientific consensus on smoking and lung cancer was solidifying. The technique was later applied by groups skeptical of climate science, by industries facing regulation on chemical exposures, by anti-vaccine activists, and by various others. The pattern is recognizable: a small number of credentialed dissenters are amplified by media and political actors as if their views represented a substantial minority of the field, even when the actual minority is tiny. Citizens reading reports of “scientific controversy” should ask whether the controversy exists at the level of the working scientific community or only in public discourse.

The Naomi Oreskes test

The historian Naomi Oreskes, in her work on climate-science discourse and on the broader history of manufactured doubt, suggests a useful test: examine where the experts publish, who they publish with, what their training is, and how their views are received in the peer-reviewed literature, not on op-ed pages or talk shows. A handful of credentialed contrarians publishing in non-specialist outlets while the rest of the field publishes in peer-reviewed journals is not a symmetric controversy. The signal of a real controversy is disagreement within the peer-reviewed literature, between researchers actively publishing in the relevant area, and engaging with one another’s arguments.

The bottom line

Scientific consensus is a sociological fact about the community of researchers — it is not unanimity, not infallibility, and not proof, but it is the considered judgment of the people who study a question for a living. It forms through convergent evidence, mechanistic understanding, replication, and formal synthesis. It can change, slowly and conservatively, as evidence accumulates. Citizens should weight consensus based on the strength of the underlying processes — high for well-established questions, more provisional for emerging ones — and should be alert to manufactured controversy in which the appearance of scientific dispute does not reflect actual scientific dispute.

What to read or watch next

  • Naomi Oreskes, Why Trust Science? (Princeton University Press, 2019). Philosophical and historical case for what makes scientific consensus trustworthy and where its limits are.
  • Naomi Oreskes and Erik Conway, Merchants of Doubt (Bloomsbury, 2010). Detailed history of how a small number of scientists, often the same individuals, helped manufacture controversy on tobacco, acid rain, ozone depletion, and climate change.
  • Intergovernmental Panel on Climate Change, Synthesis Report (2023). Example of formal scientific synthesis on a major question; the Summary for Policymakers is accessible.
  • Steven Shapin, Never Pure: Historical Studies of Science as if It Was Produced by People with Bodies, Situated in Time, Space, Culture, and Society, and Struggling for Credibility and Authority (Johns Hopkins, 2010). Useful for understanding how scientific authority is socially produced.

CHAPTER 14

Examples: Settled and Contested

It can help to have a working sense of where, in the broad landscape of public-science questions, the evidence is genuinely settled, where it is provisional, and where it remains contested. Citizens reading any individual paper benefit from the broader context: a finding consistent with a settled body of work is on different ground from a finding that contradicts one. This chapter offers a deliberately mixed selection of examples — questions ordinary citizens encounter — sorted by where the evidence currently stands. The point is not to teach the answers but to model how to think about the strength of different kinds of evidence.

Essentially settled

Some questions, despite occasional public controversy, are scientifically about as settled as anything in their domains. The evidence is overwhelming, has been replicated across many methods and decades, and the working community of relevant specialists has converged. Reasonable disagreement at the margins is possible but the central claims are not in doubt.

  • The earth is approximately 4.5 billion years old. Established by radiometric dating using multiple independent isotope systems, consistent with stellar and cosmological evidence.
  • Living species are descended from common ancestors through the process of evolution. Established by paleontological, comparative anatomical, biogeographical, and (now) genetic and molecular evidence converging across nearly two centuries of work.
  • The earth’s climate is warming, and human activities (especially the burning of fossil fuels) are the dominant cause. Established by temperature records, ice cores, ocean heat measurements, atmospheric chemistry, and physical models, with formal synthesis through the IPCC and major scientific academies.
  • Cigarette smoking causes lung cancer and cardiovascular disease. Established by decades of epidemiological work supplemented by laboratory understanding of the mechanism.
  • Vaccines (including the MMR vaccine) do not cause autism. The original 1998 Wakefield paper claiming the link was retracted in 2010 after being shown to be fraudulent, and subsequent studies of millions of children have found no link.
  • HIV causes AIDS. Established by epidemiology, virology, and the dramatic effectiveness of antiretroviral therapy in suppressing the virus and preventing progression.
  • The germ theory of disease. Microorganisms cause specific infectious diseases, established over the late nineteenth century and reinforced by the entirety of modern microbiology.

Strongly supported but with active research at the edges

Other questions have well-established central conclusions but with active research on details, mechanisms, or boundary conditions. The basic answer is not in serious doubt, but specific magnitudes, contexts, or applications remain under investigation.

  • Lead exposure in childhood has lasting cognitive effects. The basic finding is well-established; the precise dose-response curve and the threshold (if any) below which lead is safe remain under investigation.
  • Air pollution has health effects. The link to respiratory and cardiovascular disease is well-supported; the precise contributions of different pollutants and the effects at lower exposure levels are still being refined.
  • Antidepressants are effective for moderate-to-severe depression. The aggregate evidence supports moderate efficacy; the questions of magnitude, mechanism, and which patients benefit most remain debated.
  • Early childhood adversity affects later outcomes. The link is well-established; the relative contributions of prenatal, infant, toddler, and later experiences and the moderating effects of resilience factors are still under study.

Contested at the substantive level

Other questions are genuinely contested in the working scientific literature. Reasonable researchers using rigorous methods reach different conclusions, often because the evidence is genuinely mixed, the relevant comparisons are difficult, or the underlying constructs are themselves contested.

  • The optimal level of dietary salt for population health. Some research finds substantial benefit from population-level salt reduction; other rigorous research finds little benefit or even harm at very low intakes.
  • The effects of routine breast cancer and prostate cancer screening on mortality. Decades of research have produced ongoing debate about whether the benefits exceed the harms of false positives, overdiagnosis, and overtreatment.
  • The effect of moderate alcohol consumption on cardiovascular health. For decades observational evidence suggested benefit, but more recent analyses controlling for confounders have largely undermined that conclusion.
  • The effects of common psychiatric interventions on long-term outcomes. Short-term efficacy is often clearer than long-term effectiveness, and methodological debates about how to measure outcomes complicate evaluation.
  • The economic effects of moderate increases in the minimum wage. As discussed in Chapter 12, this is a long-running debate where different methods sometimes produce different results.

Preliminary or speculative

Other questions of real public interest are at the preliminary stage of scientific investigation. Findings may exist, but the literature is small, the methods are still developing, replication is limited, and confident conclusions are premature.

  • Many findings about the gut microbiome and human health. The basic biology of the microbiome is real; many specific claims about its connection to depression, autism, autoimmune disease, and behavior are at preliminary research stages.
  • Many claims about “nootropic” supplements and cognitive enhancement. Some have modest evidence; many do not.
  • Many short-term interventions claiming to durably change personality, character, or values. As Chapter 11 noted, social-science interventions are often shown to have small, context-dependent effects when rigorously studied.
  • Many specific claims about how social media affects mental health. There is real evidence of effects, but specific causal claims about specific outcomes are still developing.

How to use these examples

These lists are illustrative, not exhaustive, and reasonable people may quibble with where particular questions are placed. The point is the categories themselves: a citizen asked to evaluate a finding can usefully ask what category it falls in. A claim contradicting an essentially settled question requires extraordinary evidence; a claim refining a contested question requires only careful evidence; a claim from a preliminary area requires patience to see how it holds up. Treating every claim as if it had the same evidentiary status — either dismissing all claims as unreliable or accepting all claims as established — is a failure to distinguish among the genuine differences in how confident we are entitled to be.

The bottom line

Some scientific questions are essentially settled. Others are well-supported with active research at the edges. Others are genuinely contested. Others are preliminary or speculative. The same finding can be a discovery in a preliminary area, a refinement in a settled one, or a major revision in a contested one. Citizens reading research benefit from asking where, in the spectrum from settled to speculative, a particular claim falls. A finding’s strength is partly a function of how it stands in the broader landscape of evidence.

What to read or watch next

  • Naomi Oreskes, Why Trust Science? (Princeton, 2019). Returns here for the framework on evaluating where consensus deserves trust.
  • Cochrane Library (cochrane.org/cochrane-reviews). Searchable database of systematic reviews on clinical questions; useful for placing specific medical claims on the settled–contested spectrum.
  • National Academies of Sciences, Engineering, and Medicine (nationalacademies.org). The U.S. body that produces consensus reports on contested scientific and policy questions; reports are downloadable free.
  • Megan McArdle, “How Public Health Took Part in Its Own Downfall,” The Atlantic (May 1, 2023). On the importance of acknowledging genuine uncertainty in public-health communication.

CHAPTER 15

Reading Science Journalism

For most citizens, most of the time, research findings arrive not through the original papers but through journalism. A study is published; a press release is issued; a reporter writes an article; the article is shared on social media; commentary follows. By the time a citizen encounters a finding, it has passed through several translations — each of which can clarify or distort. Reading science journalism well is a distinct skill from reading the underlying research, and most of the time it is the more practically important skill, because it is the layer where citizens actually engage. This chapter examines the science journalism pipeline and what to watch for.

The press-release problem

A great deal of science journalism is downstream of the university press release. Press offices write announcements describing their researchers’ papers in attention-grabbing terms; reporters under deadline pressure rewrite the press releases into news stories; the stories often retain the framing of the press release. Studies of this pipeline have found systematic distortions: causal language is added where the underlying paper warranted only correlational claims; effect sizes are dropped or rounded up; caveats are omitted; the relevance to humans is asserted where the underlying research was on mice. A 2014 study in the British Medical Journal traced exaggeration in health news back primarily to the press releases, not the journalists — and found that university press offices, not corporate ones, were among the worst offenders.

The implication for citizens is that the news article is often a lightly edited version of the press release, which is in turn a tendentious version of the paper. Whenever a science news story matters to you, the press release is usually findable on the institution’s website, and the original paper is often accessible through preprint servers or institutional repositories. Comparing the news story to the press release to the abstract to the paper is illuminating. It also shows you, in microcosm, how science discourse becomes detached from science.

Common forms of distortion

Several patterns recur in science journalism, and citizens benefit from being able to recognize them:

  • Single-study sensationalism. A single new paper is reported as if it overturns a settled question or establishes a new one, with no engagement with the broader literature. Citizens should ask: how does this fit with prior work? Is this the first study or one of dozens?
  • Mouse-to-human leaps. A finding from animal models is reported as if it had been demonstrated in humans. A drug that cures a disease in mice has perhaps a 10–20 percent chance of even reaching human trials successfully. Mouse findings are early signals, not human findings.
  • Correlation reported as causation. A correlational finding is reported with causal language, often through the simple device of using “linked to” or “associated with” in the body but causal verbs in the headline.
  • Effect-size erasure. A statistically significant finding is reported as if the effect were large or important, when the actual effect size may be tiny.
  • Cherry-picked experts. The article quotes a single researcher — often the paper’s author — endorsing the interpretation, without seeking critical voices from the field.
  • Overreach about real-world implications. A study of a specific population in a specific context is described as having implications for everyone everywhere.

The diet-and-cancer pattern

A particularly notorious genre is the diet-and-disease story. Almost every food has, at some point, been reported as either causing or preventing cancer. Dr. Jonathan Schoenfeld and Dr. John Ioannidis examined 50 randomly selected ingredients from a cookbook and found that 80 percent had been the subject of published research linking them to cancer risk — sometimes as protective, sometimes as harmful, sometimes both. The underlying problem is the combination of small effect sizes, observational data, multiple comparisons, and publication bias — conditions that produce, on average, a steady stream of “newly discovered” associations that mostly do not survive scrutiny. Citizens encountering a diet-and-disease headline should be deeply skeptical, and should look for the meta-analytic synthesis (often by Cochrane or by groups like the World Cancer Research Fund) rather than the latest single study.

Good science journalism

Distinguished science journalism does exist and is one of the most useful forms of public communication. Its hallmarks are: engagement with the underlying methodology rather than just the headline finding; explicit reporting of effect sizes in concrete terms; quotation of independent experts including critics; placement of the new finding in the context of the existing literature; honest description of remaining uncertainty. Outlets and reporters who consistently meet this standard are scarce but worth following — The New York Times’s science section, The Atlantic’s science writing, the science pages of The Economist, in-depth pieces in Wired and The Verge, science podcasts like Radiolab, and individual reporters like Carl Zimmer, Ed Yong, Apoorva Mandavilli, and others have produced consistently strong work. The Knight Science Journalism Program at MIT and the Science Communication Program at UC Santa Cruz train many of the practitioners. Following a few high-quality science journalists can be more valuable than reading any number of press-release rewrites.

Social media and viral findings

Social media has changed how scientific findings reach the public, mostly for the worse. A graph stripped of context, a single sentence from an abstract, an extracted quote can travel in hours to millions of people, often without the original paper visible at all. The findings that go most viral are typically the ones that confirm pre-existing partisan beliefs or that strike emotional chords; whether they are correct rarely correlates with whether they spread. Citizens engaged with science via social media should treat any finding that arrives stripped of citation as preliminary at best and probably inaccurate. The cost of looking up the underlying source is real but small; the cost of forming beliefs from misunderstood viral content is much larger.

The 2020s information environment

The COVID-19 pandemic was a stress test of science communication and revealed many problems. Findings from preprints — not yet peer-reviewed — went viral and shaped policy before they could be vetted. Reasonable scientific uncertainty, present at every stage of the pandemic, was at times communicated as confident certainty, undermining trust when the picture later changed. Findings inconvenient to particular political positions were sometimes ignored or dismissed by all sides. The lessons for citizens are not new but were thrown into relief: hold initial reports loosely, expect updating, distinguish genuine controversy from partisan posturing, and acknowledge genuine uncertainty when it exists.

The bottom line

Most citizens encounter research through journalism, which routinely amplifies, simplifies, and distorts the underlying findings. Press releases shape news stories; news stories add causal language; social media strips context. Common distortions include single-study sensationalism, mouse-to-human leaps, correlation reported as causation, and cherry-picked experts. Good science journalism exists — careful, skeptical, contextual — but is rare. Citizens benefit from following a few high-quality science journalists, comparing news stories to original papers when stakes are high, and treating viral findings with deep skepticism until verified.

What to read or watch next

  • Petroc Sumner et al., “The Association Between Exaggeration in Health-Related Science News and Academic Press Releases,” BMJ 349 (2014): g7015. The empirical study tracing health-news exaggeration to press releases.
  • Carl Zimmer’s science writing. Books and journalism for The New York Times set a high standard for explanatory rigor and intellectual honesty in science reporting.
  • Ed Yong’s journalism for The Atlantic. Particularly his pre-pandemic and pandemic-era work, including “The Plague Year” (Atlantic, 2020) and the body of pandemic reporting that won him a Pulitzer.
  • Health News Review (healthnewsreview.org). Now archived but still searchable; for years rated U.S. health stories on a standardized rubric. The archive is a useful tutorial in what to look for.

PART SIX

Putting It Together

A practical toolkit for citizens, an honest accounting of what research can and cannot do, and why any of this matters for self-government

CHAPTER 16

A Citizen’s Toolkit for Reading Research

The previous chapters covered a great deal of territory — study design, statistics, replication, the texture of different fields, the formation of consensus, and the distortions of science journalism. This chapter compresses what matters most into a working toolkit. The goal is not to turn readers into part-time peer reviewers but to give them a small set of habits and questions that, applied honestly, will catch most of the errors and exaggerations that mislead the public. None of this requires advanced statistics. Most of it requires the willingness to slow down and ask one or two more questions before believing or repeating a claim.

Five questions to ask of any research claim

When a study or finding crosses the reader’s path — in a news article, a social media post, a politician’s speech, an advertisement — these five questions, asked in order, will sort most claims into roughly the right bucket. None of them require reading the original paper. All of them can be answered, at least roughly, in a few minutes.

  • What was actually measured, and on whom? Was the study in mice, in cells, or in humans? If humans, who — college students, hospital patients, a representative national sample, an online convenience sample? How many people, and over how long? A claim about “what causes Alzheimer’s” based on twelve mice fed an unusual diet for six weeks is a different kind of finding from one based on twenty thousand humans followed for thirty years. The press release will rarely make this distinction; the study almost always does.
  • Was it an experiment or an observation? Did researchers assign the treatment, or did they observe what people already do? If observational, what reasonable confounders might explain the result instead of the proposed cause? Most public-facing research claims of the form “X causes Y” rest on observational studies that cannot, by themselves, establish causation. Demand a higher standard of evidence — randomized trials, multiple study designs, or strong natural experiments — before accepting a causal claim.
  • How big is the effect, in real terms? Strip away percentage changes and look at absolute numbers. A 50% increase in a one-in-a-million risk is still a one-in-half-a-million risk. A medication that reduces heart attacks from 3% to 2% is helpful, but the risk reduction is one percentage point, not the more impressive-sounding “33% relative reduction.” Effects that sound dramatic in headlines often shrink to modest size on inspection.
  • Has it replicated? Is this the first study to find this result, or the tenth? A finding consistent across many studies, by independent teams, with different samples and methods, is far more credible than a single splashy result. New, surprising findings deserve interest — and skepticism. Wait for replication before reorganizing one’s life around them.
  • Who did it, and who paid? Researchers at reputable universities and research hospitals are not infallible, but they have professional incentives toward honesty and reputational stakes in being right. Industry-funded research can still be valid — most pharmaceutical research is industry-funded — but readers should know about funding and look for independent replication. Studies from advocacy organizations, think tanks with strong ideological positions, or in-house corporate research deserve extra scrutiny, not because they are necessarily wrong but because the incentives are mixed.

Red flags in coverage and claims

Beyond the questions one should ask, there are signals that a research claim is more likely to be unreliable, or that the coverage of it is more likely to be distorted. None is automatically disqualifying, but the more flags present, the more skepticism is warranted.

  • Single study. A single study, however striking, is the beginning of a conversation, not the end. Be especially wary of headlines that announce a finding as if it were settled when it has not yet been replicated.
  • Tiny sample. Studies of fewer than a few dozen subjects, particularly when claiming subtle effects, are unlikely to be reliable. Look for sample size in the article or methods section; if you cannot find it, be more skeptical.
  • Surprising effect from an obscure source. Findings that contradict large bodies of prior research deserve interest — sometimes the contrarian is right — but the burden of proof rises with the size of the claim. A single small study overturning thirty years of evidence is much more likely to be wrong than right.
  • Mouse to human in one leap. Mouse studies are useful for hypothesis generation and basic biology. They are not human studies. Coverage that suggests otherwise is misleading.
  • Causal language for observational data. Phrases like “study finds X causes Y” applied to surveys, cohort studies, or other non-experimental work are almost always overstatements. Look for hedged language in the original paper; that hedging is usually correct, and its removal is the journalist’s editorial choice.
  • Press release as primary source. If the news story’s only source is a university or company press release — and not an independent expert or the underlying paper — the coverage is likely to share the press release’s spin. Press releases are advocacy documents; they are not impartial summaries.
  • Lone-genius narrative. Stories about a single brilliant scientist whose findings are being suppressed by the establishment occasionally describe real injustice. More often they describe someone whose work cannot withstand scrutiny and who has retreated into a narrative of persecution. The default should be skepticism.
  • Round numbers and decisive verdicts. Real research findings tend to be hedged, qualified, and inconclusive. “Scientists prove X” is almost always overstatement. “Research suggests X under conditions Y, with caveats Z” is closer to how scientists actually talk.

Where to look when stakes are high

For most everyday research claims, the questions and red flags above are sufficient. For claims that matter more — a medical decision, a vote on a ballot measure, a major lifestyle change — it is worth going further. The good news is that the resources for citizens have never been better, and most of them are free.

  • Find the original paper. If a news story names a study and a journal, search for the paper. Many are open access. Read the abstract, then the discussion section, where authors typically acknowledge limitations. Even readers who cannot follow the methods can usually understand what the authors themselves think their study does and does not show.
  • Look for systematic reviews. For medical and health questions, systematic reviews from the Cochrane Collaboration aggregate the best available evidence on specific clinical questions. They are written for clinicians but are usable by motivated lay readers, and their structured “summary of findings” tables are unusually clear. For other fields, look for review articles in major journals — articles whose job is to summarize and assess a body of literature.
  • Check Retraction Watch. If you find a striking study, search the title or authors at retractionwatch.com. If the paper has been retracted, this database will say so. The site also runs original journalism on misconduct cases.
  • Use Google Scholar to see the citation trail. Searching Google Scholar (scholar.google.com) for a paper’s title shows how many times it has been cited and by whom. A study cited only by its authors and a handful of others, particularly years after publication, has not had much impact. A study cited hundreds of times by independent researchers has, at minimum, been engaged with seriously.
  • Consult expert summaries. For major scientific questions, organizations like the National Academies of Sciences, Engineering, and Medicine, the Cochrane Collaboration, the Intergovernmental Panel on Climate Change, and the Centers for Disease Control publish careful syntheses written for non-specialists. These summaries are not infallible, but they reflect the assessments of working scientists and represent the best available consensus, with appropriate caveats, on contested questions.
  • Read more than one source. If three different reputable outlets cover a finding and describe it the same way, with similar caveats, the underlying claim is probably reasonably represented. If only one outlet has the story, or if the framing varies wildly across outlets, more skepticism is warranted.

Calibrating your confidence

Reading research well is largely about calibration: matching one’s confidence in a claim to the strength of the evidence behind it. Most public discussion of research errs in one of two directions — too credulous, accepting any claim attached to the word “study,” or too dismissive, treating all research as politically motivated nonsense. Both are mistakes. The goal is graduated confidence: high confidence in well-established findings backed by many converging lines of evidence, moderate confidence in more recent but plausible findings supported by some replication, low confidence in single studies or contested fields, and very low confidence in claims that contradict broad consensus and rest on weak evidence.

Calibration also means being willing to update. Findings that seemed solid sometimes fall; findings that seemed questionable sometimes hold up. Holding views about empirical questions provisionally — strongly enough to act on, loosely enough to revise — is among the harder cognitive habits to develop. It is also among the most valuable. The opposite habits, on display constantly in public life, are settled certainty about questions where the evidence is genuinely uncertain, and reflexive skepticism about findings that are inconvenient. Both produce confident error.

The bottom line

Reading research well does not require advanced training. It requires asking five questions of any claim: what was measured and on whom; was it experimental or observational; how big is the effect in real terms; has it replicated; and who did it and who paid. It requires alertness to red flags: single small studies, mouse-to-human leaps, causal language for observational data, press release sources, lone-genius narratives, and decisive verdicts where the evidence is mixed. For high-stakes claims, citizens can find the original paper, look for systematic reviews, check for retractions, and consult expert summaries from bodies like the National Academies or Cochrane. Above all, the work of reading research well is the work of calibration — holding beliefs in proportion to evidence, and updating when evidence changes.

What to read or watch next

  • Center for Open Science’s public-facing resources (cos.io). Practical guides to reading and evaluating research, written for non-specialists.
  • Cochrane Library plain-language summaries (cochrane.org). For health questions, the Cochrane reviews include lay summaries that translate technical findings into clear prose.
  • Retraction Watch (retractionwatch.com). A working database of retractions and ongoing journalism on scientific misconduct — useful both as a reference and as a window into how peer review fails when it does.
  • Tom Chivers and David Chivers, How to Read Numbers (2021). A short, accessible guide to evaluating quantitative claims in news, written by a science journalist and an economist.

CHAPTER 17

What Research Can and Cannot Tell You

A guide to reading research would be incomplete without honest discussion of what research is for. Public discussion of science often oscillates between two errors: treating empirical research as though it could settle questions that are fundamentally about values, and dismissing research findings as mere opinion when they conflict with one’s priors. Both errors share a common root — confusion about the proper jurisdiction of empirical inquiry. This chapter draws the lines as clearly as honesty allows: what research is genuinely good at, what it can do only partially, and what it cannot do at all.

What research can tell you

Within its proper domain, research can do several things very well. Understanding these capabilities helps citizens use research where it is useful and avoid demanding from it answers it cannot supply.

  • Describe the world. Research excels at description: how many people have a condition, how a disease progresses, what students typically learn at what ages, how often a behavior occurs in a population, how a chemical reacts under specified conditions. Descriptive research, when done with adequate samples and methods, produces reliable knowledge about how things are.
  • Identify causal relationships, when conditions allow. Well-designed experiments can establish that A causes B with high confidence. Randomized controlled trials, natural experiments with strong instruments, and converging evidence from multiple study designs together build causal knowledge. Smoking causes cancer. Vaccines prevent disease. Lead poisoning impairs cognition. These are not opinions; they are the result of decades of converging causal evidence, and citizens are entitled to treat them as facts.
  • Predict, with bounded accuracy, what will happen. Engineering, climate science, epidemiology, and economics make predictions — about bridges, ice sheets, disease spread, and unemployment — that are routinely useful. Predictions come with uncertainty, and the uncertainty is often itself the most important information. But research-based prediction, properly understood, is far better than guessing.
  • Test whether interventions work. Does this drug reduce mortality? Does this curriculum improve reading? Does this policy reduce crime? These are empirical questions that research can address — not always with certainty, but with structured evidence that beats untested intuition. The evidence is often disappointing: many interventions that seem promising do not work, or work less well than hoped. This is itself useful information.
  • Map the consequences of choices. Research can illuminate the trade-offs of different paths. A given economic policy is likely to produce certain effects on employment, prices, and growth, with associated costs and benefits. Research cannot tell anyone whether the trade-offs are worth it; that is a values question. But it can tell them, with bounded confidence, what they are choosing among.
  • Reduce uncertainty over time. Even on contested questions, careful research, accumulated over decades, narrows the range of plausible answers. Where two views were possible in 1980, evidence may have ruled one out by 2020. The history of science is, at its best, the history of uncertainty progressively reduced through empirical work. This is the source of science’s genuine cultural authority — not the infallibility of any given study, but the long-run record of getting closer to truth.

What research can do only partially

Some questions sit at the boundary of research’s proper domain. Empirical findings are relevant but not decisive. Pretending research has resolved these questions when it has not is a common failure of public discourse.

  • Predict complex social systems. Research can model economies, but the financial crisis of 2008 caught most economists off guard. Research can study elections, but forecasting models often fail. Research can examine policy interventions, but transferring findings from one context to another is fraught. Citizens should expect humility from researchers studying complex social systems and should grant only modest authority to confident predictions in such domains.
  • Settle questions that depend heavily on definitions. How many “mass shootings” occur per year, how many “extreme weather events” have intensified, what the “poverty rate” is, whether “inequality” is rising — each of these depends on choices about how to define and measure. Reasonable definitions can produce different answers. Research can describe what happens under each definition; it cannot settle which definition is correct, because that question is partly conceptual.
  • Resolve questions about rare events or unprecedented situations. Research is built on patterns in data. For genuinely rare events — a particular kind of crisis, a new technology, a unique geopolitical configuration — there may not be enough comparable cases to support strong conclusions. Researchers may have informed views, but they will not have rigorous findings.
  • Predict individual outcomes from population data. Research can tell you that smokers, on average, have a higher mortality rate than non-smokers. It cannot tell you whether your specific neighbor will get cancer. Research can identify drugs that work for the average patient. It cannot promise that a particular drug will work for you. The leap from population averages to individual cases is real, but it always carries uncertainty.
  • Generalize beyond the studied population. A finding established in twenty-year-old American college students may or may not hold for fifty-year-old farmers in Vietnam. A drug tested mostly on white men may behave differently in women or in people of other ancestries. A program that worked in one school district may fail in another. The boundaries of generalization are an empirical question themselves, and they are often unsettled.

What research cannot tell you

Some questions lie outside the empirical domain entirely. No study, however well done, can answer them. Pretending otherwise is to clothe values judgments in the borrowed authority of science — a practice that is dishonest, even when one agrees with the conclusion.

  • What you should value. Research can tell you that a policy will reduce crime by 10% but cost $50 million. It cannot tell you whether crime reduction is worth that cost. That requires weighing values — safety against fiscal prudence, present spending against future benefits, the interests of one group against another. Such weighing is the job of citizens and their representatives, not of empirical inquiry.
  • What is morally right. Empirical research describes how the world is; it cannot determine how it ought to be. The fact that something is widely done, or that it has certain consequences, does not by itself determine whether it is right. Research can inform moral judgments — by clarifying consequences, identifying who is affected, and disclosing what is at stake — but it cannot replace them.
  • What the law should be. Legal questions involve tradition, precedent, constitutional structure, and democratic legitimacy alongside empirical considerations. A finding that some practice causes harm is relevant to whether it should be regulated, but the move from “causes harm” to “should be illegal” involves weighing rights, costs, alternatives, and unintended consequences. Research informs these judgments; it does not determine them.
  • Whose suffering matters more. Research can document who is affected by various policies and how. It cannot adjudicate competing claims of injury, dignity, or recognition. These are moral and political questions, and pretending that empirical findings settle them is a confusion of categories.
  • What gives life meaning. Questions of purpose, meaning, vocation, beauty, faith, and how to live a good life are not empirical. Research can describe what people report finding meaningful, what correlates with reported well-being, what choices typical lives make. It cannot tell anyone what to live for. That belongs to philosophy, religion, art, and the considered judgment of individuals — not to any study.

The trap of “the science says”

In public debate, both sides are tempted to invoke “the science” as a trump card — a way to short-circuit moral and political argument by claiming empirical authority. “The science says we must do X.” “The science is settled on Y.” Sometimes these claims are accurate — there genuinely is overwhelming evidence on smoking, vaccines, the age of the earth, and many other questions. But often “the science says” smuggles values into what sounds like a factual claim. Science describes; values prescribe. When someone says “the science says we must adopt this policy,” they are usually making a values argument disguised as an empirical one.

This trap operates across the political spectrum. Activists of all stripes claim scientific authority for conclusions that mix empirical findings with contested values. Politicians invoke studies that support their preferred policies and dismiss studies that do not. Journalists, in search of clean stories, present complex evidence as decisive verdict. The honest move — hard to make, easy to fail at — is to separate the empirical claim from the values judgment. “Research suggests this policy would have these effects, and given my values, I think we should adopt it” is honest. “The science says we must adopt this policy” is usually not.

Citizens who recognize this distinction become harder to manipulate. They can grant the empirical claim while contesting the values argument; they can disagree about what to do without denying the underlying findings. This is, in the end, what makes democratic argument possible at all. If every disagreement collapses into a fight about whether the science is real, common ground becomes impossible. If empirical questions are kept separate from moral and political ones, citizens with different values can still reason together about the world they share.

The bottom line

Research is genuinely powerful within its proper domain. It can describe the world, identify causes, test interventions, predict consequences, and reduce uncertainty over time. It is more limited at the boundary: predicting complex social systems, generalizing across contexts, and settling questions that depend on definitions or rare events. It cannot tell anyone what to value, what is morally right, what the law should be, or how to live a good life — those are questions for democratic deliberation and individual conscience. The phrase “the science says”, when used to settle a values dispute, almost always smuggles a moral argument into what sounds like an empirical one. Honest argument keeps empirical claims and values claims separate, granting each its proper authority.

What to read or watch next

  • Daniel Sarewitz, “Science and Politics: A Productive Tension,” Issues in Science and Technology (various essays). Sarewitz, a philosopher of science and policy scholar, has written extensively on the proper roles of empirical research in democratic decision-making.
  • Roger Pielke Jr., The Honest Broker: Making Sense of Science in Policy and Politics (2007). A scholar of science policy distinguishes among the roles scientists can play in public debate — pure scientist, science arbiter, issue advocate, honest broker — and what each demands.
  • Hilary Putnam, The Collapse of the Fact/Value Dichotomy (2002). A philosopher’s argument that the fact/value distinction is more porous than often thought, with implications for how research bears on moral questions.

CHAPTER 18

Why Research Literacy Matters for Democracy

This guide began with the assumption that ordinary citizens should be able to read and evaluate academic research. That assumption is contestable. One could argue that research is a specialist activity, that public understanding of it will inevitably be partial, and that the best citizens can do is defer to credentialed experts. The argument has some force; nobody can be a competent judge of every field, and humility about one’s own grasp of complex evidence is a virtue. But pure deference is also dangerous, both for the citizen and for the polity. This final chapter argues that research literacy — not expertise, but informed competence — is essential to self-government in a world increasingly shaped by what research finds and what it claims to find.

Democracy and empirical claims

Modern democratic argument is saturated with empirical claims. Should the country build more nuclear plants? The argument turns partly on safety statistics, climate impact, and economic competitiveness — all empirical questions. Should marijuana be legalized? Empirical questions about health effects, criminal justice consequences, and tax revenue intersect with moral judgments. Should this drug be approved? Should this curriculum be taught? Should this regulation be tightened or loosened? Behind almost every consequential public decision lies a thicket of empirical claims, often contested, often weakly grounded, often confidently asserted by partisans on multiple sides.

Citizens who cannot evaluate research are dependent on others to evaluate it for them. In a healthy society, this dependence is reasonable: experts in fields one does not understand are part of the division of cognitive labor that makes complex civilization possible. But the dependence becomes corrosive when the trust is misplaced — when the experts are not, in fact, doing their work well, when their authority is borrowed for purposes other than their findings, or when bad-faith actors exploit the public’s inability to distinguish good research from bad. The citizen who cannot tell a randomized trial from a survey, a single study from a meta-analysis, or a peer-reviewed paper from a press release is entirely at the mercy of whoever is loudest.

The crisis of expert authority

Public trust in scientific and expert institutions has declined over the past generation, with sharp drops following specific high-profile failures. The opioid crisis, in which authoritative medical guidance and aggressive pharmaceutical marketing combined to produce mass addiction; the financial crisis of 2008, which most economists failed to predict; aspects of the COVID-19 pandemic response, in which official guidance shifted, sometimes for legitimate reasons and sometimes not; and the steady drumbeat of high-profile retractions and fraud cases all contribute. The replication crisis, discussed at length in Part III of this guide, has further dented expert authority within fields and outside them.

The decline in trust is not entirely irrational. There are real failures in the institutions of expertise; there are real cases where credentialed authorities got important things wrong; there are real moments when scientific authority has been borrowed for political purposes. At the same time, the decline in trust has not been accompanied by improved citizen ability to evaluate evidence. The public has become more skeptical without becoming more discerning, more willing to dismiss experts without becoming better at distinguishing the experts who deserve dismissal from those who do not. The result is a polity where confident pseudoscience and confident genuine science compete on something approaching equal terms in public discussion, and where the answer to “whom should we believe?” often resolves to “whoever shares my politics.”

Restoring a healthier relationship between citizens and expertise will require work on both sides. Experts and their institutions must be more honest about uncertainty, more rigorous in their methods, and more careful about distinguishing professional findings from personal opinions and political advocacy. Citizens, for their part, must develop enough literacy to engage critically with claims rather than accepting or rejecting them based on tribe. Neither side can do its part alone, and this guide is an attempt to support the work the citizen side has to do.

How research illiteracy is exploited

Bad-faith actors of all political stripes exploit public confusion about research in predictable ways. Recognizing the patterns helps inoculate against them.

  • Single-study weaponization. Find a single study supporting one’s position, however weak; cite it as if it were definitive; ignore the larger body of conflicting evidence. This works because most readers cannot distinguish a single small study from established consensus. Used by industry shills denying harm, by activists overstating risk, and by partisans on every contested topic.
  • Manufactured doubt. On topics with strong scientific consensus — the link between smoking and cancer, the human contribution to climate change, the safety and efficacy of vaccines — well-funded actors have at various points cultivated doubt by emphasizing genuine uncertainty at the margins, sponsoring contrarian studies, and elevating fringe views to apparent equivalence with mainstream findings. The strategy succeeds because most citizens cannot weigh the relative quality of evidence and so treat “there is debate” as equivalent to “we don’t know.”
  • Conflated jurisdictions. Use empirical findings, real or fabricated, to settle questions of values. “Science proves we must do X” substitutes scientific authority for democratic argument. The technique succeeds when citizens are unable to distinguish empirical from normative claims.
  • Cherry-picked experts. On any contested question, find a credentialed expert who agrees and feature them prominently; treat the broader scientific community’s view as one opinion among many. Effective because credentials look the same on the screen even when the substance differs.
  • Numerator without denominator. Cite the absolute number of incidents — of crime, of vaccine reactions, of immigration arrests, of school shootings — without context about rates, base populations, or trends. Effective because raw numbers are vivid and many citizens are unaccustomed to converting between counts and rates.

None of these techniques is exclusively the province of any one political faction. Each has been used by industry and by activists, by the right and by the left, by candidates and by office-holders, by journalists and by online commentators. The proper response is not to identify which side does it more but to develop the personal habits that make it harder to manipulate any individual citizen — the habits this guide has tried to lay out.

Research literacy and democratic citizenship

Self-government, as the founders understood it and as it has been reaffirmed by every generation since, depends on citizens able to deliberate intelligently about public affairs. The deliberation must rest on shared facts, however contested the values to be applied to them. When citizens cannot agree on facts — not because the facts are contested but because they cannot evaluate the evidence — democratic argument breaks down. Each side comes to see the other as not merely wrong but inhabiting a different reality. The shared deliberative space, never large in a diverse society, shrinks toward zero.

The remedy is not enforced consensus or appointed experts deciding what is true. The remedy is widespread citizen capacity to evaluate evidence — not perfectly, not equally well across all domains, but well enough to spot the more obvious failures and to extend appropriate trust where appropriate. This capacity is a civic virtue. It is built, like other civic virtues, through education, practice, and exposure to good models. It is at risk, as other civic virtues are, when its supports erode.

The eighteen chapters of this guide will not, by themselves, produce a research-literate citizenry. That is the work of schools, of public institutions, of journalism that takes its educational role seriously, of universities that resist the temptation to retreat into specialization, and of citizens themselves who choose to take the work seriously. But the guide is meant as a contribution: a set of tools that any motivated reader can use, an argument that this kind of literacy is worth cultivating, and a refusal of both technocratic deference and populist dismissal as adequate stances toward the world of research.

A closing word

The reader who has reached this point has done substantial work. The chapters covered an enormous amount of ground — from the structure of a journal article through the mechanics of statistical inference, the history of replication failures, the differences among research traditions, the formation of consensus, the distortions of journalism, and the practical questions of what research is for. None of it was easy. None of it was meant to be.

What the citizen reader should now possess is not certainty — the topic does not allow it — but a working competence: the ability to read a news article about research without being either credulous or reflexively skeptical; the ability to find an original paper and form a rough judgment of it; the ability to ask the right questions of a politician or pundit citing a study; the ability to distinguish settled science from contested findings, and both from outright fraud or pseudoscience. This is enough. It will not always be enough to reach the truth on every contested question — nothing will be — but it is enough to make the citizen a participant in democratic argument rather than a target of it.

The world that produces academic research is messy, often disappointing, sometimes corrupt, occasionally heroic, and on the whole the most powerful instrument humans have devised for understanding their world. Citizens who can read it intelligently — holding it to its proper standards, expecting from it what it can give, demanding from it what it owes, and refusing what it cannot honestly provide — are citizens better equipped to govern themselves and to share that work with their fellows. That is what this guide has been for.

The bottom line

Modern democracy is saturated with empirical claims, and citizens who cannot evaluate research are entirely dependent on others to evaluate it for them. Public trust in expert institutions has declined for partly legitimate reasons, but the decline has produced more skepticism, not more discernment. Bad-faith actors across the political spectrum exploit research illiteracy through single-study weaponization, manufactured doubt, conflated jurisdictions, cherry-picked experts, and decontextualized numbers. Research literacy — not expertise, but informed competence — is a civic virtue, learnable through practice. It cannot guarantee correct answers; it can guarantee that citizens are participants in democratic argument rather than targets of it.

What to read or watch next

  • Tom Nichols, The Death of Expertise: The Campaign Against Established Knowledge and Why It Matters (2017). A polemical but readable examination of declining trust in expertise and what it costs democratic society. Worth reading both for its arguments and for the criticisms others have leveled at it.
  • Cailin O’Connor and James Owen Weatherall, The Misinformation Age: How False Beliefs Spread (2019). A philosophy-of-science treatment of how false beliefs propagate through populations — including educated ones — and what defenses are possible.
  • Naomi Oreskes, Why Trust Science? (2019). A historian and philosopher of science makes the affirmative case for trusting scientific consensus while acknowledging legitimate grounds for skepticism. Originally delivered as Tanner Lectures.
  • Yuval Levin, A Time to Build (2020). On the broader institutional erosion of which the crisis of expert authority is one part — and what rebuilding would require.

Appendix A: Glossary

Brief definitions of key terms used in this guide. These are not formal scientific definitions but plain-language explanations adequate for understanding research and journalism about it.

  • Abstract. The short summary at the beginning of a research paper, typically 150–300 words, presenting the question, methods, main findings, and conclusion. The first thing experienced readers read; often the only thing journalists read.
  • Causal inference. The process of moving from observed correlations to claims about what causes what. Requires either experimental control (random assignment) or strong reasoning about possible confounders, often through natural experiments or instrumental variables.
  • Cohort study. An observational study that follows a defined group of people forward in time, recording their exposures and outcomes. Useful for studying conditions where experimental manipulation is impossible or unethical.
  • Confidence interval. A range of values within which a true parameter (effect size, mean, etc.) is plausibly located given the data. A 95% confidence interval, on standard interpretation, is constructed so that 95% of such intervals from repeated samples would contain the true value. Often more informative than a p-value.
  • Confounding variable. A factor associated with both the proposed cause and the proposed effect, which can produce a misleading apparent relationship between them. Controlling for confounders is the central methodological challenge of observational research.
  • Convenience sample. A sample drawn from whoever is easily available — college students, hospital patients, online survey respondents — rather than from a defined population by random selection. Convenient but often unrepresentative.
  • Effect size. A measure of how large a difference or relationship is, expressed in units that allow comparison across studies. Distinct from statistical significance, which only indicates whether the effect is detectable given sample size.
  • HARKing. “Hypothesizing After the Results are Known.” Presenting a finding that emerged from exploratory analysis as though it had been predicted in advance. Inflates the apparent strength of evidence.
  • Meta-analysis. A statistical synthesis of results from multiple studies on the same question. Useful for aggregating evidence but can be distorted by publication bias and by including low-quality studies.
  • Natural experiment. A situation in which something resembling random assignment occurs by accident — a policy change in one state but not another, a lottery, a discontinuity in eligibility — allowing causal inference without a designed experiment.
  • Observational study. A study that observes what happens without intervening. Includes cohort studies, case-control studies, and cross-sectional surveys. Cannot establish causation as cleanly as experiments but is sometimes the only ethical or practical option.
  • P-hacking. The practice of analyzing data in many ways and reporting only the analyses that produce significant results. A major source of false positives in published research.
  • P-value. The probability of observing data at least as extreme as the data actually observed, assuming the null hypothesis (typically: no real effect) is true. Not the probability that the hypothesis is wrong; not the probability that the result will replicate. Widely misunderstood.
  • Peer review. The process by which a manuscript is evaluated by other researchers in the field before publication. Essential to academic publishing but imperfect as a quality filter — especially against fraud and against errors that require re-running the analysis to detect.
  • Power. The probability that a study, given its sample size and the true effect size, will detect a real effect if one exists. Underpowered studies (low power) miss real effects and inflate apparent effects when they do find them.
  • Preregistration. Filing a study’s hypotheses, methods, and analysis plan in a public registry before data collection. Distinguishes confirmatory from exploratory analysis and prevents many forms of p-hacking and HARKing.
  • Publication bias. The tendency for studies with statistically significant or surprising results to be published while studies with null or expected results sit in file drawers. Distorts the apparent strength of evidence in any field.
  • Randomized controlled trial (RCT). An experiment in which subjects are randomly assigned to treatment or control conditions. The gold standard for causal inference, especially in medicine, but limited in social and economic questions where random assignment is impossible or unethical.
  • Replication. Repeating a study — by the original or by independent researchers — to see whether its findings hold. Direct replications use the same methods; conceptual replications test the same idea with different methods. Both have value.
  • Retraction. The formal withdrawal of a published paper, usually because of fraud or serious error. Tracked at retractionwatch.com.
  • Statistical significance. A finding is “statistically significant” if its p-value falls below a chosen threshold (usually 0.05). Does not mean the effect is large, important, or certain to replicate.
  • Systematic review. A structured synthesis of all studies meeting defined criteria on a specific question. May include a meta-analysis. Generally more reliable than any single study.
  • Type I and Type II errors. Type I: a false positive (concluding there is an effect when there is not). Type II: a false negative (missing a real effect). The choice of significance threshold trades these off.
  • WEIRD samples. Acronym from Henrich, Heine, and Norenzayan (2010): Western, Educated, Industrialized, Rich, Democratic. Refers to the over-reliance of psychology and other social sciences on samples drawn from a narrow slice of humanity, raising questions about how broadly findings generalize.

Appendix B: Quick-Reference Resources

Practical resources for citizens who want to engage further with research. All links and organizations are free or have free public-facing components unless otherwise noted. Inclusion does not imply endorsement of every position taken; these resources are useful, with critical reading, even where one disagrees with particular conclusions.

For finding research papers

  • Google Scholar (scholar.google.com). A free academic search engine. Find papers by title, author, or topic; see citation counts and trace which later papers have cited a given paper. Indispensable starting point.
  • PubMed (pubmed.ncbi.nlm.nih.gov). The U.S. National Library of Medicine’s database of biomedical research. Free abstracts; many papers fully free. The standard tool for medical and health research.
  • SSRN (ssrn.com). The Social Science Research Network. Free working papers and preprints in economics, law, finance, and other social sciences.
  • arXiv (arxiv.org). Free preprint server for physics, mathematics, computer science, and related fields. Papers here are not peer-reviewed but are widely used in those communities.
  • NBER Working Papers (nber.org). The National Bureau of Economic Research distributes working papers from leading economists, often before formal publication. Subscription for full PDFs; abstracts free.

For evaluating research quality

  • Cochrane Library (cochrane.org). Systematic reviews on medical and health interventions. Plain-language summaries available.
  • Retraction Watch (retractionwatch.com). Database of retracted papers and journalism on scientific misconduct. Searchable by author, journal, and topic.
  • PubPeer (pubpeer.com). A platform for post-publication peer review. Comments on papers from researchers who have spotted potential problems. Useful for sanity-checking high-profile findings.
  • OSF (osf.io). The Open Science Framework. Hosts preregistrations, datasets, and study materials. Useful for checking whether a study’s analysis was preregistered.

For science journalism and synthesis

  • STAT News (statnews.com). Reuters-style coverage of medicine, biotech, and health. Generally rigorous; some content is paywalled.
  • Nature News and Science News. The news sections of the two leading general-science journals. Written for scientists but readable; cover developments inside research communities that mainstream outlets miss.
  • Quanta Magazine (quantamagazine.org). Long-form science journalism on mathematics, physics, biology, and computer science. Free, edited at a high standard.
  • FiveThirtyEight’s science archives. While the site’s focus has shifted, its archived work on science statistics, replication, and methodology remains useful.

For learning more about how research works

  • Center for Open Science (cos.io). Resources on preregistration, replication, and open science practices. Public-facing materials suitable for non-specialists.
  • Sense About Science (senseaboutscience.org). A UK-based charity that produces public guides to topics like peer review, statistics, and how to read research.
  • National Academies of Sciences, Engineering, and Medicine (nationalacademies.org). Reports on contested empirical questions, written for educated lay readers. Free PDFs of all reports.

Books worth owning

For citizens who want a small library on research literacy, a few books provide especially solid foundations. None requires statistical training; all reward careful reading.

  • Stuart Ritchie, Science Fictions: How Fraud, Bias, Negligence, and Hype Undermine the Search for Truth (2020) — the most accessible single-volume introduction to the replication crisis.
  • Tom Chivers and David Chivers, How to Read Numbers (2021) — short, practical guide to evaluating quantitative claims.
  • Daniel Kahneman, Thinking, Fast and Slow (2011) — the cognitive psychology of judgment under uncertainty, including thoughtful self-criticism in later editions about which findings have replicated.
  • Hans Rosling, Factfulness (2018) — less about reading research per se, more about the larger habit of seeing the world clearly through data.
  • Jordan Ellenberg, How Not to Be Wrong (2014) — a working mathematician’s tour of statistical reasoning, with applications from voting to medicine.
  • Richard McElreath, Statistical Rethinking (2nd ed., 2020) — for the more committed reader, the most humane introduction to Bayesian statistics now in print.

Appendix C: References and Further Reading

Selected references for the major studies, books, and reports cited in this guide. Arranged by chapter for ease of follow-up. Citations follow a hybrid Chicago style adequate for finding the source. Many of the journal articles are available open access; others may require library access.

On the structure and process of science

Thomas S. Kuhn. The Structure of Scientific Revolutions. 4th ed. Chicago: University of Chicago Press, 2012 (orig. 1962).

Karl Popper. The Logic of Scientific Discovery. London: Routledge, 2002 (orig. 1934).

Robert K. Merton. “The Normative Structure of Science.” In The Sociology of Science: Theoretical and Empirical Investigations. Chicago: University of Chicago Press, 1973.

Richard P. Feynman. “Cargo Cult Science.” Caltech Commencement Address, 1974.

On peer review and publication

Drummond Rennie. “Guarding the Guardians: Research on Peer Review.” JAMA 256, no. 17 (1986): 2391–2392.

Richard Smith. “Peer Review: A Flawed Process at the Heart of Science and Journals.” Journal of the Royal Society of Medicine 99, no. 4 (2006): 178–182.

Adam Mastroianni. “The Rise and Fall of Peer Review.” Experimental History (Substack), December 13, 2022.

On p-values and statistical practice

Ronald L. Wasserstein and Nicole A. Lazar. “The ASA Statement on p-Values: Context, Process, and Purpose.” The American Statistician 70, no. 2 (2016): 129–133.

Valentin Amrhein, Sander Greenland, and Blake McShane. “Scientists Rise Up Against Statistical Significance.” Nature 567 (2019): 305–307.

Steven Goodman. “A Dirty Dozen: Twelve P-Value Misconceptions.” Seminars in Hematology 45, no. 3 (2008): 135–140.

On the replication crisis

John P. A. Ioannidis. “Why Most Published Research Findings Are False.” PLOS Medicine 2, no. 8 (2005): e124.

Open Science Collaboration. “Estimating the Reproducibility of Psychological Science.” Science 349, no. 6251 (2015): aac4716.

C. Glenn Begley and Lee M. Ellis. “Raise Standards for Preclinical Cancer Research.” Nature 483 (2012): 531–533.

Joseph P. Simmons, Leif D. Nelson, and Uri Simonsohn. “False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant.” Psychological Science 22, no. 11 (2011): 1359–1366.

Andrew Gelman and Eric Loken. “The Garden of Forking Paths.” Working paper, Department of Statistics, Columbia University, 2013.

Daryl J. Bem. “Feeling the Future: Experimental Evidence for Anomalous Retroactive Influences on Cognition and Affect.” Journal of Personality and Social Psychology 100, no. 3 (2011): 407–425.

Stuart Ritchie. Science Fictions: How Fraud, Bias, Negligence, and Hype Undermine the Search for Truth. New York: Metropolitan Books, 2020.

Timothy M. Errington et al. “Investigating the Replicability of Preclinical Cancer Biology.” eLife 10 (2021): e71601.

On samples and generalization

Joseph Henrich, Steven J. Heine, and Ara Norenzayan. “The Weirdest People in the World?” Behavioral and Brain Sciences 33, no. 2–3 (2010): 61–83.

On medical and clinical research

Marcia Angell. The Truth About the Drug Companies: How They Deceive Us and What to Do About It. New York: Random House, 2004.

Ben Goldacre. Bad Pharma: How Drug Companies Mislead Doctors and Harm Patients. London: Fourth Estate, 2012.

FDA. Guidance for Industry: E9 Statistical Principles for Clinical Trials. Washington, DC: U.S. Food and Drug Administration, 1998.

On economics, education, and policy research

David Card and Alan B. Krueger. “Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania.” American Economic Review 84, no. 4 (1994): 772–793.

Joshua D. Angrist and Jörn-Steffen Pischke. Mostly Harmless Econometrics: An Empiricist’s Companion. Princeton: Princeton University Press, 2009.

Esther Duflo and Abhijit Banerjee. Poor Economics: A Radical Rethinking of the Way to Fight Global Poverty. New York: PublicAffairs, 2011.

Robert J. Slavin. “Evidence-Based Reform in Education: What Will It Take?” European Educational Research Journal 7, no. 1 (2008): 124–128.

On science journalism and the public

Petroc Sumner et al. “The Association Between Exaggeration in Health-Related Science News and Academic Press Releases: Retrospective Observational Study.” BMJ 349 (2014): g7015.

Carl Zimmer. Various works for The New York Times and books, including Life’s Edge: The Search for What It Means to Be Alive. New York: Dutton, 2021.

Ed Yong. “How the Pandemic Defeated America.” The Atlantic, September 2020.

On expertise, science, and democracy

Tom Nichols. The Death of Expertise: The Campaign Against Established Knowledge and Why It Matters. New York: Oxford University Press, 2017.

Naomi Oreskes. Why Trust Science? Princeton: Princeton University Press, 2019.

Cailin O’Connor and James Owen Weatherall. The Misinformation Age: How False Beliefs Spread. New Haven: Yale University Press, 2019.

Roger A. Pielke, Jr. The Honest Broker: Making Sense of Science in Policy and Politics. Cambridge: Cambridge University Press, 2007.

Daniel Sarewitz. “How Science Makes Environmental Controversies Worse.” Environmental Science & Policy 7, no. 5 (2004): 385–403.

Philip E. Tetlock and Dan Gardner. Superforecasting: The Art and Science of Prediction. New York: Crown, 2015.

Yuval Levin. A Time to Build: From Family and Community to Congress and the Campus, How Recommitting to Our Institutions Can Revive the American Dream. New York: Basic Books, 2020.

On scientific fraud and misconduct

Uri Simonsohn, Joseph Simmons, and Leif Nelson. Data Colada (datacolada.org), particularly the series of posts on the Francesca Gino case beginning June 2023.

Adam Marcus and Ivan Oransky. Retraction Watch (retractionwatch.com). Various ongoing reporting on misconduct, retractions, and corrections in scientific publishing.

William Broad and Nicholas Wade. Betrayers of the Truth: Fraud and Deceit in the Halls of Science. New York: Simon & Schuster, 1982. Older but still useful overview of historic and structural issues.

————

This guide was prepared as a citizen’s reference. Errors in summary or interpretation are the responsibility of the present compilation, not of the cited authors. Citizens are encouraged to consult the original sources, which are richer, better hedged, and often more interesting than any summary can convey.

Ask this guide

Ask a question and get an answer drawn only from Reading Research. It won't make things up — if the answer isn't here, it'll tell you.