Two numbers decide whether this call needs speed or care. The first is the bill you pay to correct it once it turns out wrong. The second is how much of the error an extra week of work actually removes. Price both and the setting picks itself. Price neither and you will run your own temperament on every decision the company makes.
Reversibility is a different question and a narrower one. It asks whether the door swings back. The correction bill asks what walking back costs, how long you sit on the wrong side before anyone notices, and what burns while you sit there. A reversible decision with a six-month detection lag will hurt you more than an irreversible one you can price today.
The second number is the one almost nobody tracks. Call it the yield of care: error removed per extra day spent. In some environments it is high. In others it sits near zero, and the extra days buy confidence instead of accuracy.
Why these three models
This piece runs the Wire Model. Score the features of the decision, route to a small ensemble of formal models, then force the ensemble to produce dated actions. The scores that carried the routing:
- Historical-analog density (0.8). Clinical judgment, software change control, hospital bed planning and East African digital lending already ran this experiment at scale.
- Cognitive and affective distortion (0.8). More information reliably raises confidence. It does not reliably raise accuracy.
- Threshold and congestion (0.75). Care is a claim on a scarce reviewer, and reviewer load has an elbow past which delay stops behaving linearly.
- Regime-break risk (0.6). Machine-written analysis is collapsing the price of looking careful, which moves both sides of the trade at once.
- Deep uncertainty (0.55). You cannot observe the outcome of the decision you did not make.
That routes to three: comparative statics for the dial itself, threshold and congestion for what happens when everyone chooses care, and the luck-skill continuum for whether care can pay at all in your market. Equilibrium, complex, random. Three outcome types, so the errors point in different directions.
Behavioral and governance are folded rather than shipped as separate cards. The behavioral result is a parameter set wrong inside the loss function the first model already describes, so it earns a paragraph there and no lever of its own. Governance is the third card: Mauboussin’s continuum is load-bearing here, so it gets promoted.
The framework: the setting is the correction bill divided by the yield of care
1. The dial: which way each variable pushes
Every decision carries two costs. Expected damage, the chance of being wrong multiplied by the correction bill. And delay cost, which accrues every day the decision stays open. Care buys down the first and adds to the second. The optimum sits where one more day of work removes less damage than it costs in delay.
The correction bill has three parts and only one of them is the fix. Detection lag, the time you run wrong before you find out. Burn rate, what the wrong state costs per day. Repair cost, the work to unwind it. Detection lag multiplies the other two, which is why it deserves the attention they get.
The NIST-commissioned study of software testing infrastructure puts illustrative numbers on the escalation. A requirements-stage defect costs one unit to fix when caught in the same stage, five at coding, ten at integration and system test, fifteen at beta, and thirty after release.9 The error did not change. The date you found it did.
Translate that to an operator. A price change on one WhatsApp order channel is visible in a day, so the correction bill is one day of margin and a broadcast message. A mispriced twelve-month distributor contract surfaces at renewal, so the bill is a year of leaked margin plus a renegotiation you enter from behind. Same decision type. Two different settings.
Trace each variable and the direction rule falls out:
- Correction bill rises, move toward care.
- Detection lag rises, move toward care, or spend the time on instrumentation rather than analysis.
- Decision frequency rises, move toward speed. The overhead of care is paid on every instance and compounds.
- Yield of care falls, move toward speed regardless of the stakes.
The last one is the rule operators refuse to apply: when extra work removes no error, high stakes are an argument for deciding faster, not slower. High stakes plus low yield means the deliberation is theatre and the delay is real.
Here is the folded behavioral layer. Firms do not compute these variables. Across more than 1,200 managers surveyed at global companies, fewer than half said decisions were timely and 61% said at least half the time spent making them was ineffective, priced at roughly 530,000 days of manager time a year at a typical large firm. The same work separates big-bet, cross-cutting, delegated and ad hoc decisions, and finds firms running one process weight across all four.1 Temperament is not a fourth model. It is one parameter set wrong and then defended as culture.
Comparative statics, the dial lens
Assumes: care and speed trade against each other on a smooth curve with an interior optimum.
Fits because: historical-analog density scored 0.8, so the parameters are estimable from other people’s data.
Breaks when: speed and quality co-move rather than trade, which happens whenever the binding constraint is batch size or feedback latency instead of thinking time.
Counteracts: the belief that being careful is always the responsible choice.
May reinforce: false precision, and treating a guessed correction bill as a measured one.
2. The queue: care is rationed capacity, and it has an elbow
Choosing care is not a private act. It puts the decision in a line in front of a reviewer who is serving every other decision too. Queueing theory is specific about what follows. As that reviewer’s utilization rises, average delay rises at an increasing rate. The curve has an elbow, past which small increases in load produce dramatic increases in delay, and delay approaches infinity as utilization approaches one. Higher variability drags the elbow left.5
One founder reviewing everything is a single server with high service-time variability. At moderate load a decision waits a day. Add three more a week and it waits three weeks. The curve gives no warning before the elbow, which is why teams read the collapse as a culture problem rather than a capacity one.
A review that adds two days of thinking and eleven days of waiting sold you two days of quality at the price of thirteen.
The evidence that the gate often buys nothing is direct. Organizations requiring an external body, a change advisory board or a senior manager, to approve significant changes were 2.6 times more likely to be low performers, and the researchers found no evidence that the formal approval was associated with lower change failure rates.6 The gate was installed to buy quality. It delivered delay and left the quality exactly where it found it.
So measure the wait, not the work. For your last twenty decisions, log the date the question was raised and the date it was answered. That gap, minus the hours anyone actually spent, is what your care setting costs. Nobody has ever billed you for it.
Threshold and congestion, the capacity lens
Assumes: decisions arrive at a reviewer with finite throughput, and waiting is a real cost.
Fits because: the threshold feature scored 0.75, and the delay response to load is nonlinear rather than proportional.
Breaks when: the queue is not the constraint, for example when the delay is external and the reviewer is idle.
Counteracts: costing a review by the hours booked instead of the days elapsed.
May reinforce: stripping out review that was genuinely load-bearing, because the queue argument applies to every gate equally.
3. The ceiling: care only pays where the environment is readable
The dial assumes extra work removes error. That holds only in some places. Activities sit on a continuum from all skill to all luck, and the position governs how much a small sample tells you. Where skill dominates, a few observations reveal the answer. Where luck carries a large share, you need many observations before signal separates from noise.4 Your market sits somewhere on that line and you probably cannot say where.
The conditions for trusting a fast judgment are stated precisely. Skilled intuition develops only where the environment is regular enough to be predictable and the person has had adequate opportunity to learn those regularities through practice and feedback. Fail either condition and the feeling of confidence carries no information about accuracy.3
What the extra information does instead is measured. Thirty-two judges, eight of them clinical psychologists, read a case in four instalments and answered the same twenty-five questions after each. Accuracy did not increase significantly across the stages. Confidence increased steadily and significantly, and every judge but two finished overconfident.2 The yield of care hit zero while the invoice kept running.
The reverse holds too. In a study of eight firms in a fast-moving industry, the teams that decided fastest used more information than the slow teams, not less. The difference was the kind: real-time operating and competitive data rather than forecasts. Fast teams also developed more alternatives and weighed them simultaneously instead of in sequence, and that pattern tracked better firm performance.8 Much of what feels like care is latency in your reporting loop wearing a suit.
East African digital lenders chose the extreme fast setting on purpose. The product is instant, automated and remote, with credit decisions made from airtime and handset data in seconds.10 Phone surveys of 3,150 Kenyans and 4,574 Tanzanians found 47% of Kenyan and 56% of Tanzanian digital borrowers had repaid late, with 12% and 31% reporting a default.10 Those rates are not a failure of the model. They are priced into the interest rate. Speed is safe exactly where you find out fast, and nowhere else. Those lenders learn inside thirty days. Copy the setting only if your feedback arrives that quickly.
Luck-skill continuum, the ceiling lens
Assumes: outcomes mix skill and chance in a ratio that is stable enough to estimate.
Fits because: deep uncertainty scored 0.55, and the yield of care is unobservable without a position on this line.
Breaks when: your sample is too small to locate yourself, which for an early company it always is.
Counteracts: the assumption that more diligence is always more rigour.
May reinforce: fatalism, and using luck as cover for a decision process that was genuinely sloppy.
GEER: the levers, cheapest first
Four channels carry the exposure: the correction bill, the detection lag, the yield of care, and the load on your review capacity. Reach for the cheap reversible levers before the structural ones.
- Score your last twenty decisions. Three columns: correction bill in money, detection lag in days, days held open. Costs an hour and usually ends the argument.
- Publish one threshold. Any decision under a stated correction bill gets a named owner and a 24-hour clock, no meeting. Hits queue load immediately.
- Instrument before you deliberate. Name the metric and the read date for every fast decision. Cutting detection lag is usually cheaper than cutting the error rate.
- Cap what one reviewer holds. Six open decisions, not sixteen. That moves you back down the curve instead of up it.
- Shrink the bill instead of buying accuracy. One region, one branch, a 50-unit LPO before the 500-unit one, a two-week price test on a single WhatsApp channel. Same information, smaller invoice if wrong.
- Measure the yield of care. Run two lanes for a month on one class of decision, one fast and one studied. Compare at day 60. Costs weeks, settles the question permanently.
No-lever flag: where the correction bill is large and the detection lag is long and the yield of care is unmeasured, no process setting rescues you. That is a research problem, not a decision problem. Buy information first and put the decision back in the queue.
RADAR: the portfolio, dated
DO NOW, by T+3 days. Reversible, and correct under every scenario.
- Build the twenty-decision ledger. Correction bill, detection lag, days open.
- Publish the delegation threshold as a number, not a principle.
- Put a named owner and a decision date on every open question. Anything without both is not a decision, it is a mood.
- Kill one standing review that nobody can name a reversal from.
HEDGE, by T+14. Cheap insurance against having the dial set wrong.
- Instrument the three decisions with the longest detection lag. A metric and a read date each.
- Start the two-lane comparison on one class of recurring decision.
- Write a one-page reversal log. Every decision you undo, with what it cost. This is how you get a measured correction bill instead of a guessed one.
DEFER AND TRIGGER. Irreversible and expensive to correct, so wait, but pre-commit the observable now.
- Defer: exclusive distributor terms, a core banking or ledger migration, a senior hire with equity, any contract denominated in a currency you do not earn.
- Trigger to commit: the studied lane beats the fast lane on the same decision class by T+45. That is your evidence the yield of care is positive. Move those calls to the slow setting.
- Counter-trigger: the two lanes are indistinguishable at T+45. Raise the threshold, delegate the class, and spend the freed reviewer capacity on detection.
For the person holding the queue. DO NOW: state, in currency, the correction bill above which you want to be consulted. HEDGE: report your queue depth and median days-to-answer weekly, since the wait you impose is invisible to you and only to you. DEFER: any new approval step until you can name three reversals the last one caught.
CHAIN: what history says happens next
The right comparison set is not other founders. It is any organisation that answered a bad outcome by routing a high-frequency decision through an approval body. Software change management, hospital bed planning and clinical case judgment share that structure while sharing no surface at all.
Their base rate is consistent. The gate adds waiting, and the quality it was bought to deliver never shows up in the failure data.6, 5, 2 Second order, the queue lengthens, work batches up while it waits, and it arrives in larger, riskier lumps. Third order, those larger batches raise the failure rate, which reads as proof that even more review is needed. The loop closes and tightens on itself.
Adjust for the present state before you copy anyone. A careful, successful team may just be operating in a forgiving market with short detection lags, and a fast, lucky one may be selecting the easy decisions to make fast. Strip both effects out and most of the observed gap between careful firms and quick ones goes with them. Detection lag survives the adjustment, which is why it is the variable worth funding.
Matrix-break flag. Care is becoming cheap to produce, and that breaks the model’s central assumption. Machine-written analysis costs close to nothing, so the deliberation side of the trade collapses in price while looking identical from outside. The delivery data shows the other half: for every 25% increase in AI adoption, an estimated 1.5% reduction in delivery throughput and a 7.2% reduction in delivery stability, attributed to weak fundamentals like batch size and testing rather than to the tools.7 The correction bill rises while the price of looking careful falls. Both parameters move at once, in opposite directions. Recalibrate your threshold on measured bills within two quarters. The old one was set in a world where analysis was expensive.
What this ensemble cannot see
Four gaps, and none of them close with more thinking.
The yield of care on a single decision. You never observe the version of the call you did not make. The two-lane comparison is a workaround, and it only works on decisions that repeat.
Your position on the luck-skill line. Locating it needs a sample your company does not have, and the small sample you do have is the one most likely to mislead you.
Correction bills nobody has paid. Every estimate in your ledger was written by someone who has not yet unwound the thing they are pricing. Expect the real bills to run high.
Reputational damage. A visible error with a distributor, a regulator or a lead investor costs more than its repair line, and it arrives as a jump rather than a slope. No model here produces that number.
None of that changes Monday. Write the twenty-decision ledger, publish one threshold in currency, and put a read date on the three decisions whose errors would surface last. All of it by T+3. The dial you cannot set from first principles you can set from your own reversal log by T+45.
Sources and notes
- De Smet, A., Jost, G., and Weiss, L. “Three keys to faster, better decisions.” McKinsey Quarterly, May 2019. Survey of more than 1,200 managers; fewer than half say decisions are timely, 61% say at least half the time spent making them is ineffective, roughly 530,000 manager-days a year at a typical Fortune 500 firm. The decision typology (big-bet, cross-cutting, delegated, ad hoc) is from the same piece. McKinsey’s own domain refuses automated requests, so this links a full PDF copy of the article hosted by PAISBOA.
- Oskamp, S. “Overconfidence in Case-Study Judgments.” Journal of Consulting Psychology 29(3), 1965, 261-265. Thirty-two judges, eight of them clinical psychologists, four successive stages of case information, a twenty-five-item test at each stage. Full text.
- Kahneman, D., and Klein, G. “Conditions for Intuitive Expertise: A Failure to Disagree.” American Psychologist 64(6), 2009, 515-526. The two necessary conditions are a high-validity (sufficiently regular) environment and an adequate opportunity to learn its regularities. Full text.
- Mauboussin, M. J. “Untangling Skill and Luck.” Legg Mason Capital Management, 15 July 2010. The skill-luck continuum and the sample-size argument are in the opening section and the section on factors shifting activities toward luck. Full text.
- Green, L. V. “Queueing Theory and Modeling.” Columbia Business School. Delays rise at an increasing rate with utilization, the curve has an elbow, average delay approaches infinity as utilization approaches one, and higher variability moves the elbow left. Full text.
- Forsgren, N., Smith, D., Humble, J., and Frazelle, J. “Accelerate: State of DevOps 2019.” DORA and Google Cloud. Respondents were 2.6 times more likely to be low performers where an external body approved significant changes, and the authors found no evidence that formal approval was associated with lower change failure rates. Report PDF, heavyweight change process section.
- “Accelerate State of DevOps Report 2024.” DORA and Google Cloud. An estimated 1.5% reduction in delivery throughput and 7.2% reduction in delivery stability for every 25% increase in AI adoption, Figure 10. Report PDF, with the landing page at dora.dev.
- Eisenhardt, K. M. “Making Fast Strategic Decisions in High-Velocity Environments.” Academy of Management Journal 32(3), 1989, 543-576. Inductive study of eight microcomputer firms; fast decision makers used more, not less, information, favoured real-time operating data over forecasts, and considered multiple alternatives simultaneously. Full text.
- RTI International for the National Institute of Standards and Technology. “The Economic Impacts of Inadequate Infrastructure for Software Testing,” Planning Report 02-3, May 2002. The one / five / ten / fifteen / thirty cost-factor progression is Table 5-1, published there as an illustrative example rather than a measured average; the $22.2 billion to $59.5 billion annual estimate is from the executive summary. Report PDF.
- Izaguirre, J. C., Kaffenberger, M., and Mazer, R. “A Digital Credit Revolution: Insights from Borrowers in Kenya and Tanzania.” CGAP Working Paper, October 2018. National phone surveys of 3,150 Kenyans (1,037 digital credit users) and 4,574 Tanzanians (1,132 users), June to August 2017: 47% of Kenyan and 56% of Tanzanian digital borrowers repaid late, 12% and 31% report having defaulted. The instant, automated and remote characterisation is the paper’s own. Working paper PDF. The word in the title is the authors’, not ours.