Skip to main content

Why Nudges Fail

Jason Hreha· Updated September 5, 2026

Summary: Large-scale field programs and publication-bias analyses show why a pooled nudge estimate should not be treated as a forecast for a new intervention. Effects depend on the intervention, population, outcome, and context. Small positive effects can still matter when costs are low. These findings do not demonstrate that Behavioral Strategy outperforms nudging. BS-0003 BS-0027

This page summarizes what the evidence implies about “nudge-first” work and links to the deeper analyses.

Summary (the decision rule) #

  • Use relevant evidence: favor interventions, populations, and outcomes close to the proposed application; do not assign every new nudge the same expected effect. BS-0027

  • Compare benefits and costs: a small absolute lift may be worthwhile at low cost, but may not meet a project’s required outcome. The cited research does not determine that decision for every setting. BS-0003

  • Specify the outcome for defaults: enrollment, repeated action, and downstream outcomes are different measures. A default can affect one without establishing the others. See the default analysis.
  • Cautionary tale: “defaults did it” narratives (e.g., organ donation) often misattribute outcomes to a checkbox change rather than infrastructure and process. BS-0004

What we mean by “nudge” on this page #

By “nudge” we mean choice architecture interventions (defaults, framing, reminders, simplification) intended to shift behavior without changing the underlying value proposition or materially changing incentives.

Two clarifications:

  • Definitions vary. Simplifying a form can alter choice architecture and practical feasibility at the same time. Describe the actual intervention rather than relying only on the label.
  • Nudges can be ethical and transparent; the critique here is primarily effect size, reliability, and strategic leverage.

What the best evidence says #

Note: Some sources summarize nudge outcomes as percentage-point lifts in specific programs, while others report standardized effect sizes (Cohen’s d) across many studies. These metrics are not directly comparable; taken together, they primarily imply small average effects and substantial heterogeneity.

1) At-scale field RCT programs: small average effects #

In the two US nudge units studied (126 RCTs, ~23M individuals), average effects are ~1.4 percentage points - about one-sixth of the ~8.7 percentage-point average in the academic-journal comparison sample. BS-0003

The nudge-unit average was positive and statistically significant. Its practical value depends on costs, scale, and the outcome. The comparison is between nudge-unit and academic-journal samples; it should not be described as a simple laboratory-to-field failure.

2) Publication-bias correction: pooled effects collapse toward ~0 #

Bias-correction work argues that publication bias is severe in the nudge literature and that once you correct for it, the mean effect moves toward zero. BS-0027

A 2025 second-order meta-analysis (14 meta-analyses; 1,638 primary studies; ~30M participants) reports an aggregated effect (d = 0.270) that drops to ~0 (d = 0.004) after publication-bias adjustment, while noting that many underlying meta-analyses are low quality. BS-0027

3) Baseline meta-analyses: bigger averages, weak forecasting value #

A broad meta-analysis reports average effects around d ≈ 0.43-0.45, with substantial heterogeneity and evidence of publication bias. Treat this as a descriptive average - not a reliable forecast for a new nudge in a new context. BS-0011


What these findings mean for a project #

A practical question is:

Does evidence for this intervention and context justify a test, given the required effect, cost, and alternatives?

When evidence is weak, reduce the scale of the initial commitment and compare plausible options. This site’s practitioner recommendation is to examine:

  • selecting a behavior with strong Dispositional/Capability/Context Fit,
  • changing feasibility (tools, infrastructure, workflow),
  • building repeatable value and feedback loops.

That recommendation is not a comparative finding from the nudge literature. No study cited here tests whether Four-Fit or DRIVE outperforms a nudge-led process. See the evidence status.


Defaults: distinguish the setting from the outcome #

Defaults can be useful when the target “behavior” is really a one-time configuration decision (e.g., auto-enrollment) or when the environment can ethically set a recommended option with easy opt-out.

Even in “canonical success” domains like retirement savings, participation gains do not automatically translate into large long-run wealth effects once turnover and withdrawals are accounted for. BS-0055

A default can maintain a setting without creating a deliberate routine. That distinction does not make the downstream outcome unimportant. Measure the outcome the program is intended to produce and its duration, rather than treating either enrollment or habit formation as sufficient proof. See the default analysis.


Organ donation: a cautionary case #

Organ donation is the headline default story, but outcomes depend on the system: donor identification, ICU pathways, trained coordinators, logistics, governance, and family conversations, far beyond legal default status.

See: Organ Donation Defaults and BS-0004 .


Practical guidance (how to talk about nudges credibly) #

When testing a nudge:

  1. Specify the behavior, denominator, and window.
  2. Pre-commit to a minimum effect size that justifies the effort and any ethical cost.
  3. Plan rollback if effects are null or if the intervention harms trust.
  4. Report null, negative, and positive results, with uncertainty and the relevant costs.

Behavioral Strategy proposes a fit-first sequence. Choice architecture can be one candidate intervention within that sequence; its value must be measured in the intended setting.