A premortem works by taking the doubt out of the question. Ask a team what could go wrong and you have asked them to rank possibilities, so they answer in categories: execution risk, market risk, key-person risk. Tell them the company is dead and give them the date, and you have asked them to explain a fact. A fact gets answered with episodes.
That switch is the whole method. The certainty does the work. The time travel is decoration.
The second effect is smaller to describe and larger in the room. Naming the failure stops being an act of disloyalty and becomes the assigned task.
Why these three models
This piece runs the Wire Model: score the features of the decision, route to a small ensemble of formal models, then force the ensemble to produce dated actions. The scores that mattered:
- Cognitive and affective distortion (0.85). The binding constraint is not what your team knows. It is what the question lets them retrieve.
- Many-agent aggregation under social influence (0.8). The output is a group estimate. Whether it is any good is decided in the first two minutes.
- Strategic actors with conflicting interests (0.7). The person holding the killing objection also holds a job, a promotion path and a relationship with the person who wrote the plan.
- Deep uncertainty (0.75). You cannot enumerate the failure space, so you cannot compute over it.
- Historical-analog density (0.8). Startup and megaproject post-mortems are abundant and the causes repeat.
- Regime-break risk (0.4). Human retrieval is stable machinery. What is changing is the cost of producing a list of failure modes, which is now close to zero.
That routes to three models: behavioral (how certainty changes retrieval), mechanism design (how the rules change what gets said), and crowd aggregation (how the room’s answers combine, or fail to). They span three outcome types, complex, equilibrium and random, so their errors point in different directions.
Two notes on what is folded in. The governance layer sits inside the mechanism design card, because on this topic governance is the mechanism: who speaks, in what order, and who may not answer back. And the base-rate work is not a fourth card because the structure already carries it. CHAIN is the base-rate section.
The framework: a retrieval device, a rule change and an aggregator
1. The prompt: certainty is the active ingredient
Gary Klein put the premortem into circulation in 2007, and cited a 1989 result: prospective hindsight, imagining that an event has already occurred, increases the ability to correctly identify reasons for future outcomes by 30%.1 That number is the reason most people run the exercise. It is also the reason most people run it wrong.
Go to the 1989 study itself. Mitchell, Russo and Pennington varied two things: temporal perspective, whether the event sits in the future or the past, and uncertainty, whether the event is certain or merely possible. Their finding was that temporal perspective showed little influence. Uncertainty did the work. Explanations for sure events were longer, contained a higher proportion of episodic reasons, and were expressed in the past tense.2
Doubt makes the mind answer in categories. Certainty makes it answer in episodes. An episode has a date, an actor and a sequence. A category has none of those, which is why a risk register full of categories never changes a decision.
So the wording carries the effect. Write the failure as a fact with numbers attached, not as a possibility. Weak version: imagine the launch does not go well. Working version: it is eighteen months from today, we shut the company, last full month of revenue was 2.1 million shillings against a 9 million plan, and the board meeting where we agreed to stop lasted eleven minutes.
Then let the room answer in the same register. Agents stopped restocking in week six because the margin did not clear their transport cost. The distributor signed the exclusivity and did nothing with it for fourteen months. Our float sat with a partner whose settlement file we could never reconcile, so we could not prove volumes to the bank.
The effect is measurable. In a controlled comparison, participants who ran a premortem cut their confidence in the plan more than those who imagined success or did a filler task, F(2,51) = 4.02, p = .024, with an effect size of d = .81.3 That is a large effect for a ninety-minute meeting. Note what was measured: confidence and reasons. Not survival.
Behavioral, the retrieval lens
Assumes: the team already holds the relevant knowledge, and the constraint is access to it rather than possession of it.
Fits because: cognitive and affective distortion scored 0.85, and the documented mechanism is uncertainty, not time.
Breaks when: the failure mode is genuinely outside the room’s experience, a new regulation, a currency move, a competitor nobody has met. Retrieval cannot reach what was never encoded.
Counteracts: the belief that a longer risk list is a better one.
May reinforce: vividness. The most cinematic failure story wins attention it has not earned.
2. The rules: make the disloyal sentence someone’s job
Mechanism design asks one question. What rules make self-interest produce the outcome you want? Here the outcome you want is a specific, costly truth from a person who pays for saying it.
Under normal meeting rules, that person pays twice. They look like the one who is not on board, and they get handed the problem they named. Klein is blunt about it: projects fail partly because people are reluctant to raise reservations during planning, and what surfaces in a premortem is the kind of thing participants ordinarily would not mention for fear of being impolitic.1
Confidence in the team does not fix this. Across 51 work teams in a manufacturing company, psychological safety predicted learning behaviour, and team efficacy did not once safety was controlled for.4 A team can be certain it will win and still say nothing. Those are separate variables and founders routinely read the first as evidence of the second.
The usual fix is also wrong. Brainstorming’s famous rule is do not criticise. Tested against instructions that actively encouraged debate and criticism, in the United States and France, the debate condition produced significantly more ideas than minimal instructions, while the no-criticism rule did not differ significantly from giving no instructions at all.5 Banning criticism does not buy you safety. It buys you silence with better manners.
Design the rules. The mood follows. Five that cost nothing:
- The decision owner writes first and reads last. Order of speaking is the strongest lever in the room and it is free.
- The owner may not respond during the read-out. Not to correct, not to clarify. Questions come after every cause is on the wall.
- The person who names a cause is not assigned to fix it. Otherwise you have taxed the truth and you will collect less of it.
- Every cause gets a named owner, a date and one observable that would tell you it is happening. No observable, no entry.
- One participant has no upside in the decision. A board observer, an advisor, the finance lead of a sister company.
Mechanism design, the rules lens
Assumes: people withhold rationally, so changing the payoff changes what they say.
Fits because: strategic actors with conflicting interests scored 0.7, and the withheld information is exactly the information you need.
Breaks when: the fear is of the job, not the meeting. Where dissent has been punished before, no meeting rule reverses it in ninety minutes.
Counteracts: the assumption that an open-door culture is a substitute for a protocol.
May reinforce: proceduralism. A perfectly run session that changes nothing is still theatre.
3. The room: your yield is diversity, and you can destroy it in two minutes
A premortem is an aggregation device. What an aggregator returns depends on how different its inputs are and on whether they were formed independently.
On difference: modelling functionally diverse problem solvers, Hong and Page found that a team of randomly selected agents outperforms a team of the best-performing agents, because as the pool grows the best performers necessarily become similar to each other in the space of problem solvers.6 Read that as a staffing rule. Your four smartest people are your four most similar people. They wrote the plan together and they will fail in the same direction.
On independence: in an experiment with 144 subjects, mild social influence, simply seeing what others had estimated, narrowed the diversity of opinions without improving collective error. The authors named three separate damages: a social influence effect that shrinks diversity for no accuracy gain, a range reduction effect that pushes the truth to the edge of the group’s range, and a confidence effect that raises certainty after convergence with no improvement in accuracy.7
The first confident voice in the room does not improve the estimate. It shrinks the range and raises everyone’s confidence in a worse answer.
Klein’s original protocol already handles this: everyone writes independently, first, before anything is said out loud.1 Almost every team skips it and goes straight to discussion, which converts the exercise into a conversation about whatever the loudest person said in minute two.
Practical composition for an African operating team: the field or agent-network lead, the person who reconciles settlement, one customer-facing seller, one finance person, one outsider with no stake. Not the four people who built the model.
Crowd aggregation, the combination lens
Assumes: individual errors are partly independent, so they partly cancel when combined.
Fits because: many-agent aggregation under social influence scored 0.8, and the failure mode is documented and fast.
Breaks when: everyone shares the same blind spot. Independent inputs from identically informed people cancel nothing.
Counteracts: seniority weighting, and the habit of running the session as a discussion.
May reinforce: false comfort in headcount. Ten people in a room is not ten independent reads.
GEER: the levers, cheapest first
Four channels carry the effect: retrieval (how the prompt is worded), voice cost (what saying it costs the speaker), diversity (who is present and how inputs combine), and commitment (how much of the plan is already irreversible). Pull the cheap, reversible levers first.
- Rewrite the prompt as a dated fact with numbers. Twenty minutes of your own time. Largest single gain available.
- Seven minutes of silent writing before anyone speaks. Free. Protects the only thing the aggregator runs on.
- Owner writes first, reads last, does not defend. Free, and costs only ego.
- Add two people who did not build the plan, one with no upside. Costs a calendar invite and some discomfort.
- Convert every cause into owner, date and observable. Half a day. A cause with no observable is a feeling.
- Resize the irreversible part of the plan. Shorter contract, staged tranche, pilot instead of rollout. Costs weeks and real negotiating capital.
No-lever flag: if the irreversible commitment is already signed, a premortem produces regret, not a decision. Run it on the exit clauses instead: what would have to be true for us to use the break clause, and by when.
RADAR: the portfolio, dated
DO NOW, by T+3 days. Reversible and dominant across scenarios.
- Name the single irreversible line in your current plan. The lease, the tranche, the exclusivity, the senior hire, the licence application.
- Book ninety minutes before that line gets signed. Write the death statement yourself: date, revenue number, one sentence on the board meeting.
- Run it. Seven minutes silent, then read-out, owner last, no defending.
- Take the top five causes. Each gets an owner, a date and one observable a stranger could check.
HEDGE, by T+14. Cheap insurance on the causes you cannot instrument.
- For the two causes with no observable, buy optionality instead of certainty: a three-month contract with an extension option in place of twelve, one corridor in place of three.
- Hold back a slice of any facility you have negotiated. Undrawn capital is a hedge; drawn capital is a covenant.
- Give one outsider standing permission to send you the killing question in writing, once a quarter.
DEFER AND TRIGGER. Irreversible, so wait, and pre-commit the trigger now.
- Defer: drawing the debt. This matters more each year. Debt reached US$1.6B across 107 transactions in African tech in 2025, up 63% year on year, and made up 41% of all capital deployed.10 Debt does not forgive a wrong assumption the way equity does.
- Trigger to proceed: the top two premortem causes each have an observable that has stayed clean for 28 days. Then draw, sign, hire.
- Counter-trigger: either observable moves the wrong way by T+28. Do not draw. Re-run the ninety minutes with the two new people who saw it first.
If you are the decision owner. DO NOW: publish the rule that you speak last, in writing, before the first session. HEDGE: run one premortem a quarter on the plan you are most confident about, not the one you are most worried about. DEFER: changing your planning process until you have three sessions’ causes checked against what actually happened.
CHAIN: what usually happens next
Match the reference class on structure, not surface. The structure here is an irreversible commitment, made by a small group with shared information, under a leader who has already committed publicly, on the basis of a forecast.
Large infrastructure fits that structure exactly, and the base rate is unkind. Cost forecast inaccuracy averages 44.7% for rail, 33.8% for bridges and tunnels and 20.4% for roads. Rail passenger forecasts are off by an average of -51.4%, with 84% of rail projects wrong by more than 20% in either direction. Across the seventy years for which data exist, accuracy has not improved.8 Better models did not fix it, because the problem was never the model.
Venture-backed failure fits it too, and the causes are boring. Across 431 VC-backed companies that shut down since 2023, with 385 categorised, 70% ran out of capital, 43% had poor product-market fit, 29% cited timing or macro conditions and 19% had unsustainable unit economics.9
Present-state modifier: with debt at 41% of capital deployed on the continent,10 “ran out of capital” now arrives with a covenant and a lender attached, which shortens the window in which you can still change your mind.
Subtract the counterfactual before you over-credit the method. What the premortem is shown to do is lower confidence and surface reasons.3 No study cited here shows it raises survival. Price it as a cheap, evidenced improvement to a single high-stakes decision, and stop selling it as a weekly ritual.
Matrix-break flag. A model can now generate forty plausible failure modes in seconds. That collapses the cost of the list, which was never the scarce input. The scarce input is the private knowledge sitting with the person who visits the agents. Short run, generic risk registers get longer and worth less. Medium run, the session’s value moves entirely to ranking, ownership and observables. Design for the medium run: let a machine write the generic list before the meeting, and spend the ninety minutes on what only your people know.
What this ensemble cannot see
Three things, and the first is the largest.
What never made the list. Shown a fault tree for why a car will not start with major branches deleted, people barely noticed. The residual “all other problems” category should have risen roughly sixfold, from .078 to .468, to absorb what was missing. It only doubled. Of 55 subjects who saw a pruned tree, one weighted the residual correctly.11 A premortem builds a fault tree, quickly, from memory. You will never feel the weight of the branch that was left off the tree.
Plausibility is not probability. The ensemble ranks by how well a story was told, and a vivid, well-narrated failure will be over-insured while the ordinary one on the CB Insights list quietly kills the company.
Your own base rate. There is no dataset for your corridor, your regulator or your distributor. The reference classes above are structural analogies, and they will be wrong in specific ways you cannot see from inside.
Which leaves one action that survives all three. Name the irreversible line in your plan today. Put ninety minutes in the diary before it gets signed, write the death statement with a date and a revenue number, and require an observable against every cause the room produces. If you cannot write the observable, you are not ready to sign.
Sources and notes
- Klein, G. “Performing a Project Premortem.” Harvard Business Review, September 2007, p. 18. Source of the 30% figure as Klein reports it, the protocol instruction that participants write independently before speaking, and the “fear of being impolitic” observation. Publisher page; full text of the printed article mirrored at Southeastern Oklahoma State University.
- Mitchell, D. J., Russo, J. E., and Pennington, N. “Back to the future: Temporal perspective in the explanation of events.” Journal of Behavioral Decision Making 2(1), 1989, 25-38. DOI 10.1002/bdm.3960020103. The publisher page bot-blocks; the author abstract quoted here (temporal perspective showed little influence, uncertainty strongly affected explanation type, sure events produced longer explanations with a higher proportion of episodic reasons in the past tense) is the version deposited with Crossref and displayed at CoLab and in the Crossref metadata record.
- Keysor, E., Wojtyna, A., and Veinott, E. S. “Premortem: Evaluating two structured analytic techniques for group brainstorming.” Society for Judgment and Decision Making, 2020. Experiment 1, n = 53: F(2,51) = 4.02, p = .024, d = .81. Experiment 2 found no significant difference between individual and group premortems, F(1,42) = 0.85. The poster also reports the earlier finding of Veinott, Klein and Wiggins (2010) that a group premortem outperformed a pro-and-con list or a critique. Poster.
- Edmondson, A. “Psychological Safety and Learning Behavior in Work Teams.” Administrative Science Quarterly 44(2), 1999, 350-383. Study of 51 work teams in a manufacturing company. Full text.
- Nemeth, C. J., Personnaz, M., Personnaz, B., and Goncalo, J. A. “The liberating role of conflict in group creativity: A study in two countries.” European Journal of Social Psychology 34(4), 2004, 365-374. Debate versus minimal instructions on group ideas, F(1,33) = 5.23, p < .03; brainstorming’s no-criticism rule versus minimal, F(1,33) = 2.28, not significant. Author manuscript, UC eScholarship.
- Hong, L., and Page, S. E. “Groups of diverse problem solvers can outperform groups of high-ability problem solvers.” Proceedings of the National Academy of Sciences 101(46), 2004, 16385-16389. PNAS bot-blocks; open-access full text at PubMed Central.
- Lorenz, J., Rahwan, T., Helbing, D., and Schweitzer, F. “How social influence can undermine the wisdom of crowd effect.” Proceedings of the National Academy of Sciences 108(22), 2011, 9020-9025. Experiment with 144 subjects; social influence, range reduction and confidence effects. Open-access full text at PubMed Central.
- Flyvbjerg, B. “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal 37(3), 2006, 5-15. Cost inaccuracy of 44.7% rail, 33.8% bridges and tunnels, 20.4% roads; rail passenger forecasts averaging -51.4% with 84% wrong by more than 20%. Author preprint: arXiv 1302.3642.
- CB Insights, “The Top 20 Reasons Startups Fail,” published March 2026. Analysis of 431 VC-backed companies that shut down since 2023, of which 385 carried enough detail to categorise. Percentages sum above 100 because companies cited multiple causes. Report.
- Partech, 2025 Africa Tech Venture Capital Report. Debt of US$1.6B across 107 transactions, up 63% year on year, representing 41% of all capital deployed; US$4.1B total across 570 transactions. Release and report page.
- Fischhoff, B., Slovic, P., and Lichtenstein, S. “Fault Trees: Sensitivity of Estimated Failure Probabilities to Problem Representation.” Journal of Experimental Psychology: Human Perception and Performance 4(2), 1978, 330-344. Experiment 1: for the first pruned tree, the residual “other” category should have increased from .078 to .468 and only doubled; one subject of the 55 across both pruned groups assigned the residual an appropriate weight. Original technical report, Decision Research, August 1977: full text.