FounderWiseDecisions, not feeds
← All articles FounderWise · Long-form

The nudge you were sold does not work

Adjusted for publication bias, the average nudge effect is 0.04 with a confidence interval that includes zero. Change the payoff instead.

04 Sep 2026 12 min read By Joshua Pi’Rwot
Share X LinkedIn

Someone has proposed changing a default, reordering your pricing tiers, or rewriting a confirmation screen, and the business case rests on nudge research. The published meta-analysis put the average effect at a Cohen’s d of 0.43, a small-to-medium effect across behavioural domains.

Adjusted for publication bias, the same dataset gives 0.04, with a 95% interval of 0.00 to 0.14. The interval includes zero. Do not fund the framing change. Change the payoff.

Why these three models

The decision is whether to spend engineering time on an intervention whose evidence base has just moved. The features that fire are a claimed behavioural effect, a test you could run yourself, and a literature with a known selection problem.

Three lenses. Better-response behaviour produces a complex answer about what actually moves a choice. The value of information produces an equilibrium answer about whether your own test is worth running. The reference class produces a cycle answer about which findings survive adjustment in general, which is the one that generalises past this example.

1. What actually moves a choice: the payoff gap, not the frame

The first lens is the constructive half of this article, and it explains both why nudges underperform and what to do instead.

People do not best-respond. They better-respond: they play strategies with higher expected payoffs more often, rather than picking the best one with certainty. The responsiveness parameter governs how sharply behaviour tracks payoffs. When the payoffs facing someone are close together, behaviour approaches a coin flip no matter how the options are arranged, and when the payoffs are far apart, behaviour tracks them closely whatever the presentation.

That gives a clean diagnostic. If your two options differ by an amount the user can barely feel, no amount of reordering, defaulting or rewording will reliably move them, because there is nothing for the responsiveness to bite on. The framing is being asked to do work that only a payoff difference can do.

So the honest replacement for a nudge is not a better nudge. It is a bigger gap: a real price difference, a removed step that costs the user genuine time, an outcome they can feel. Those are more expensive than a framing change, which is exactly why the framing change was proposed.

This also explains a pattern that otherwise looks like bad luck. The interventions that do show durable effects in the field tend to be the ones that changed something material, an automatic enrolment that moves money by default, a removed form that saved a week. Those get filed under choice architecture because the change was presented as a design decision, but the work was done by the payoff, not the presentation. When a nudge works, check whether it was a nudge.

Quantal response, the payoff-gap lens

  • Assumes: choice probability rises with payoff difference rather than jumping to the maximum.
  • Fits because: the intervention under discussion changes presentation and leaves payoffs untouched.
  • Breaks when: the payoff gap is large and the failure is genuinely one of attention or comprehension, where presentation does matter.
  • Evidence: grade B. Fits experimental data well, and is criticised as flexible enough to fit a great deal.
  • Counteracts: the belief that presentation is a substitute for a real difference.
  • May reinforce: dismissing genuine interface failures as unfixable.

2. Whether your own test is worth running

The second lens is the one that stops this article becoming a licence to do nothing.

Information has value only when it can change an action. Before running any test, ask what you would do differently under each result. If a positive result and a negative result lead to the same decision, the test has zero value regardless of how well it is designed, and the effort belongs elsewhere. That is a boundary question rather than a design question: what you have drawn inside the experiment decides what it can possibly tell you.3

The framing matters because most product teams treat testing as automatically virtuous. It is not. A test consumes traffic, engineering time and, more expensively, the credibility you spend when you ask people to wait for a result. Those costs are real whether or not the test was ever going to change anything.

Applied here, the question is sharp. If the nudge is cheap and reversible and you would ship it on a positive result and drop it on a negative one, the test has value and you should run it on your own users rather than argue about someone else’s meta-analysis. If you would ship it anyway because a senior person likes it, the test has no value and you should say so out loud instead of buying evidence you will ignore.

This also sets the bar for how much of your own traffic to spend. An effect of the size the adjusted estimate suggests would need a sample most companies cannot reach. Deciding you cannot detect it is a legitimate and cheap conclusion, and it is more honest than an underpowered test that returns noise you then interpret.

Value of information, the decision-relevance lens

  • Assumes: a test is worth its cost only if some result would change what you do.
  • Fits because: the disputed evidence invites you to gather your own.
  • Breaks when: the value is political rather than decisional, and the test exists to settle an argument between people.
  • Evidence: grade A. A formal result, not an empirical claim.
  • Counteracts: testing things whose outcome you will overrule.
  • May reinforce: refusing to measure anything whose result is uncomfortable.

3. The reference class: what survives adjustment

The third lens is the one worth carrying past this example, because the specific claim will be replaced and the pattern will not.

Here is the case in full, and both numbers sit side by side in one table. Mertens and colleagues published a meta-analysis of choice-architecture interventions in which the random-effects estimate across all studies is 0.43, with a 95% interval of 0.38 to 0.48. Maier and colleagues re-analysed that same dataset correcting for publication bias, and the adjusted combined estimate is 0.04, with an interval of 0.00 to 0.14.1 Broken out by category, the adjusted information estimate is 0.00 and the adjusted structure estimate is 0.12.1

Their conclusion is quotable and worth quoting exactly: they find strong evidence of publication bias across almost all subdomains, and conclude that “after correcting for this bias, no evidence remains that nudges are effective as tools for behaviour change.”1

The base rate this establishes is not about nudges. It is that a literature with small samples, a strong prior in favour of positive results, and low publication cost for a confirmatory finding will produce an inflated pooled estimate, and that the inflation is often the whole effect. The number to ask for is not the effect size. It is the effect size after adjustment.

Reference class, the adjustment lens

  • Assumes: findings matched on publication conditions inflate at a knowable rate.
  • Fits because: the claim in front of you comes from exactly such a literature.
  • Breaks when: the field pre-registers and reports nulls, which changes the selection process and therefore the base rate.
  • Evidence: grade A as a method, and the adjustment techniques themselves carry assumptions worth reading.
  • Counteracts: accepting a headline effect size at face value.
  • May reinforce: a blanket cynicism that discards findings which did survive.

The levers, cheapest first

  • Ask for the adjusted number. One question, no cost: has this effect been re-estimated correcting for publication bias, and what did it become? If nobody knows, that is your answer for now.
  • Check whether the payoffs differ at all. If your two options are near-identical in what the user gets, no presentation change will reliably move them. Fix the difference or drop the project.
  • Price the test before designing it. Name what you would do on each result. If the actions match, cancel.
  • Ship the cheap and reversible ones anyway, without the theory. A clearer confirmation screen may be worth doing because it is clearer, not because a literature promised a behaviour change. Just do not budget for the effect.
  • Reserve the expensive work for payoff changes. A removed step, a real discount, a genuinely faster outcome. These cost more and they are the ones that move the responsiveness.

What to do before the next roadmap meeting

Do now, sized at one hour, effect immediate. Take the behavioural claim currently justifying something on your roadmap and find out whether it has been adjusted for publication bias. Reversible, nearly free, and dominant across every scenario about whether the intervention works.

Hedge, where the premium is the whole loss, live before the next build cycle. Add one line to your product brief template: what payoff difference does this change create? If the answer is none, the item is a presentation change and should be budgeted as one. The cost is a line in a template.

Defer and trigger, size fixed now. Do not audit every past decision that cited behavioural research. Pre-commit the trigger: the next time a roadmap item rests on a single published effect size, it gets the adjusted-number question before it gets engineering time. Decide now who has to answer it, because a question with no owner is not a gate.

What usually happens next

Run the break test first. Has a rule changed, has an actor entered or left, has a measurement become a target? Behavioural science has been changing its own rules through pre-registration and registered reports, which alters the selection process that produced the inflated estimates. A base rate drawn from the pre-registration era will differ, and that is a real change rather than a quibble.

If nothing broke, the pattern is consistent and it is not confined to nudging. A widely-cited effect gets re-examined, the adjusted estimate lands far below the headline, and the applied field keeps citing the headline for years because it is in the slide deck. The lag between correction and practice is the part that costs you money.

There is a second regularity worth carrying. The gap between a correction being published and it reaching practice is measured in years, and it is longest in the applied fields furthest from the original literature. Product and marketing teams are several citation hops from the source, so they receive the headline and almost never the re-analysis. That distance is not anyone’s fault and it is entirely predictable, which makes it something you can systematically exploit by simply asking one question earlier than your competitors do.

Subtract the counterfactual before you conclude your own nudge worked. A default change that shipped alongside a pricing change and a marketing push has no clean attribution, and the default is the cheapest of the three to credit.

What this ensemble cannot see

These three lenses can tell you that a pooled estimate is inflated. They cannot tell you that a specific intervention does not work on your specific users.

That gap is genuine and it cuts both ways. An adjusted average near zero is compatible with real effects in some contexts and negative effects in others, and the structure category retained the largest adjusted estimate of the three.2 Reading “nudges do not work” as “this will not work here” is the same error as reading the original headline as “this will work here”, with the sign flipped.

There is also a limit on the adjustment itself. Publication-bias corrections are models with assumptions, and different methods give different answers on the same data. The finding here is strong enough to change a funding decision and not strong enough to close a field.

And one property none of these models contains: proposing a nudge is often organisationally useful precisely because it is cheap and visible. Removing it without replacing that function leaves a gap that will be filled by something less measurable.

The one action that survives the ignorance: at the next roadmap meeting, ask one question of whichever item cites behavioural research. What is the adjusted effect size? If nobody in the room knows, move the item behind anything that changes a payoff, and revisit when someone does.

Who has to move

The person who needs this is whoever writes the product brief, and the brief template is where it sticks. The cheapest first test is a single added line asking what payoff difference the change creates. Items that cannot answer it are presentation changes, which is fine, as long as nobody is forecasting a behaviour change from them.

Sources and notes

  1. Maximilian Maier, František Bartoš, T. D. Stanley, David R. Shanks, Adam J. L. Harris and Eric-Jan Wagenmakers, No evidence for nudging after adjusting for publication bias, Proceedings of the National Academy of Sciences 119(31), 2022. Open-access copy: https://pmc.ncbi.nlm.nih.gov/articles/PMC9351501/. Table 1 compares unadjusted and adjusted effect size estimates and carries both figures used here: the random-effects combined estimate of 0.43 with a 95% interval of 0.38 to 0.48, and the adjusted combined estimate of 0.04 with an interval of 0.00 to 0.14, alongside the adjusted category estimates of 0.00 for information and 0.12 for structure. The paper reports strong evidence for publication bias across all subdomains apart from food when using only the most precise estimates, and the concluding sentence quoted in the text appears verbatim.
  2. Stephanie Mertens, Mario Herberz, Ulf J. J. Hahnel and Tobias Brosch, The effectiveness of nudging: A meta-analysis of choice architecture interventions across behavioral domains, Proceedings of the National Academy of Sciences 119(1), 2022, is the original meta-analysis whose dataset was re-analysed in note 1. The publisher bot-blocks automated requests, returning 403 to a declared browser user agent, so no link is given here. The 0.43 figure is taken from the comparison table in note 1 rather than from the original, and the two report it identically.
  3. John D. Sterman, Business Dynamics: Systems Thinking and Modeling for a Complex World, McGraw-Hill. The point in section 2 that a broad model boundary matters more than a great deal of detail, and that unanticipated effects are a sign of a narrow boundary rather than a property of the world, is from chapters 1 and 3. Used here for the framing that a test which cannot change a decision is a boundary problem rather than a design problem.

A note on why this article names the original authors. Publishing a large meta-analysis, sharing the data and code, and having it re-analysed by others is how the process is supposed to work, and the re-analysis explicitly thanks them for sharing well-documented data. That is a better outcome than a field which never revisits anything. The criticism here is of citing the 2022 headline in 2026, not of producing it in 2022.

And a note on one claim that is not here. An earlier draft mentioned a published correction to the original. The correction is real, but its publisher returns 403 to an automated request, so its text could not be read and confirmed. Under this series’ own standard a source counts only when the body carries the claim, so the claim was removed rather than cited to a link nobody checked.

Joshua Agonya Pi’Rwot, Founder.

Lock in your calls.

You’ve marked 0 of 5. Now choose how often you want the signals.

Step 1 · Pick your cadence

The DispatchWeekly · your Monday 5 callsFreealways

Step 2 · Where to send it

Personalize your BriefThe Brief

Tune every edition to the markets and industries you actually act on.

🔒 Unlock personalization — The Brief, $19.99/mo →
Free Dispatch forever · upgrade anytime · we never share your details.
Know where you stand?
The Dispatch tells you what changed. Knowing what to do about it is a different question, and it is the one FounderWise answers. Start with the free Traction Audit: 12 questions, about 3 minutes, scored out of 100.
Find out where you stand →

For teams, syndicates & programs

Recommended
Team
$15/seat · mo
Daily Brief for the whole team (min 3 seats).
  • Everyone on the same signal
  • Admin + shared watch-list
  • One invoice · ~25% off solo
Get Team →
Cohort Licence
$2,900/yr
Co-branded seats for one cohort, for accelerators, funds & programmes.
  • Up to 25 founder seats
  • Your logo, your cohort
  • The record of what your cohort committed to
Talk to us →
Pass the Dispatch on
Know a founder making these calls blind? Send them this week’s five — free, every Monday.

Decisions, not feeds. · Curated by Joshua Pi’Rwot · FounderWise · Free Audit · Store · parent of Business Growth Accelerator

Call committed. We’ll hold you to it.
Know where you stand →