A problem is complicated when the same input gives the same output twice and somebody who has solved it can hand you the answer. It is complex when the parts adapt to what you do, so the answer moves while you compute it. Both feel hard from the inside, which is why the classification error is common and expensive.
The error runs one direction. Operators bring expertise, sequencing and a plan to problems that only answer probes. The invoice lands later as a quarter spent, a team hired against an untested thesis, or a market that quietly reorganised around the move. Below is the test, and what changes once you run it.
Why these three models
This piece runs the Wire Model: score the features of the decision, route to a small ensemble of formal models, then make the ensemble produce dated actions. The scores that set the routing:
- Many-agent and emergent behaviour (0.85). Customers, agents, distributors and regulators each follow their own rule. You see the sum, and the sum has no author.
- Deep uncertainty (0.8). No defensible probability exists for how a market answers a product it has never seen.
- Cognitive distortion (0.7). Confidence earned in one domain travels into neighbouring domains where it earned nothing.
- Regime-break risk (0.5). Cheap instrumentation is turning unreadable environments into readable ones, which moves the boundary itself.
- Historical-analog density (0.35). Deliberately low. Scarce matching precedent is the signal this article exists to detect.
Three models follow. Comparative statics is the null: the machinery every plan runs on, whose licence conditions double as the diagnostic. Agent-based dynamics is the alternative: what a system does when the parts respond. The luck-skill continuum is the referee: how many observations before a result carries information. Their outcome types are equilibrium, complex and random. No two of them are wrong in the same way, which is the whole reason to carry three.
A behavioural layer belongs here. It is folded into the first card rather than shipped as a fourth, because its one lever, take feedback from outside your own head, already sits inside the diagnostic. The governance member is not folded. It is model three.
The framework: two kinds of hard
1. The null: a complicated problem lets you hold everything else still
Comparative statics is the operation you already perform. Move one variable, trace where the new equilibrium lands, assume the rest of the world sits still meanwhile. Every roadmap, hiring plan and pricing memo you have written runs on it, with that assumption unstated.
The assumption has published licence conditions. Kahneman and Klein call a task environment high-validity when stable relationships exist between objectively identifiable cues and subsequent events, or between cues and the outcomes of possible actions. Medicine and firefighting sit fairly high. Individual stock prices and long-term political forecasting sit at zero validity.1 Validity alone is not enough: skilled intuition also needs prolonged practice and feedback that is rapid and unequivocal, and where feedback misleads, in Hogarth’s wicked environments, wrong intuitions develop reliably and confidently.1
The paper also names the trap operators walk into most. Fractionation of skill: professionals with genuine expertise in some tasks get asked to judge areas where they have none, and neither they nor anyone watching can find the boundary.1 Your growth lead is excellent at paid acquisition and knows nothing more about agent-network economics than you.
When a problem is complicated and merely large, the fix is not a better expert. It is an outside view: forecast from the recorded outcomes of similar past projects rather than from the internals of this one.10
Comparative statics, the null lens
Assumes: a stable cue-to-outcome mapping, and that everything you are not moving stays where it is.
Fits because: it is the model already running underneath every plan, so its licence conditions double as the test.
Breaks when: the other variables are agents. They notice the move and reprice.
Counteracts: buying more analysis when the environment is the unreadable thing.
May reinforce: deference to credentials, and the folded-in behavioural trap of confidence travelling further than the skill did.
2. The generator: a complex problem hides its rule inside the aggregate
Schelling built the cleanest demonstration in 1971. Individuals with mild, non-extreme preferences about their neighbours produced heavily segregated neighbourhoods none of them had chosen. Keep his summary: systemic effects are overwhelming, there is no simple correspondence between individual incentive and collective results, and inferences about individual motives usually cannot be drawn from aggregate patterns.2
Apply that to your dashboard. The cohort curve, the WhatsApp order count, the agent float: each is an aggregate, and an aggregate does not carry the rule that made it. Read the rule off the chart and your plan rests on a guess.
Kenya ran the experiment at national scale. The interest rate capping law became operational on 14 September 2016, intended to lower the cost of credit and widen access. The Central Bank of Kenya’s own review found loan accounts fell 26.1% between October 2016 and June 2017, large banks down 27.7% and leading the rationing of small borrowers, and put the cost of shutting micro, small and medium enterprises out of credit at 0.4 percentage points of 2017 growth.11
The cap did not change the price of credit. It changed who the banks were willing to lend to. One variable moved, everything else moved with it, and the bill landed inside nine months.
The same market carries the other half of the lesson. The M-PESA pilot opened on 11 October 2005 with eight agent stores and one designed use case, microfinance loan repayment. Inside a few weeks the team was staring at transaction shapes nobody had specified and nobody could account for, so they commissioned research to explain their own screens. The answers included traders holding cash overnight because agent shops outlasted bank hours, money going in at one pilot town and out at the next, and a wife wiring bus fare to a husband who had just been robbed.7 None of it appeared in the specification. National launch followed in March 2007, seventeen months later.7 They bought information for those months, then shipped the product the users had written.
The forecast error is still running. Registered mobile money accounts in Sub-Saharan Africa came in 75% above the 2019 projection when 2024 data landed, past one billion, double the 2020 figure.8 Well-resourced analysts missed a five-year number by three quarters, because the system was complex.
Agent-based dynamics, the generator lens
Assumes: outcomes are produced by many local rules interacting, not by one global rule you can name.
Fits because: many-agent behaviour scored 0.85 and the aggregate demonstrably fails to encode the rule.
Breaks when: one actor dominates. A monopoly distributor or a single regulator collapses the system back to something you can model directly.
Counteracts: reading intent and mechanism off a chart.
May reinforce: learned helplessness, and the excuse that nothing can be planned so nothing gets committed.
3. The referee: how many observations make a result real
Mauboussin places any activity on a continuum from pure skill to pure luck, and the placement sets your sample size. Where skill decides the outcome, a small sample is revealing. Where luck plays a large role you need a big one, and extreme results revert to the mean fast.5
Product decisions sit nearer the luck end than founders assume. Only one third of ideas tested at Microsoft improved the metrics they were designed to improve, and a tester at Quicken Loans, after five years, put his own hit rate at about 33%.3
The same paper names the substitute behaviour. It is easier to generate a plan, execute against it and declare success on percent of plan delivered than to measure whether the feature moved anything.3 Percent of plan delivered is a complicated problem’s metric wearing the clothes of progress.
Trained probing moves behaviour measurably. Across four randomised control trials covering 759 firms and 11,463 data points, teaching founders to treat the idea as a testable hypothesis raised terminations and concentrated pivots: treated firms pivoted radically once or twice rather than never or repeatedly, and that pattern correlated with higher revenue.4 The discipline does not make you right more often. It makes you wrong faster and cheaper.
Luck-skill continuum, the governance lens
Assumes: the ratio of skill to luck in a domain is stable enough to set a required sample size.
Fits because: deep uncertainty scored 0.8, and every probe result needs a rule for how much to believe it.
Breaks when: the trials are not independent. One viral WhatsApp group makes ten results into one observation.
Counteracts: promoting a single good week into a strategy.
May reinforce: paralysis, and demanding statistical comfort you cannot afford.
The field guide: five reads, one hour
Take your largest open problem. Score each read 1 if the complicated answer holds, 0 if it does not.
- Repeat. Run the same input twice, a week apart. Same output within a tolerance set beforehand? Score 1. That is the licence for holding everything else still.
- Feedback. Does the result arrive quickly and unambiguously, or slowly and tangled with four other changes? Rapid and unequivocal scores 1. Slow or misleading feedback builds confident wrong instincts.1
- Reaction. Do other parties change behaviour because you acted? A price the market reprices around scores 0. A machine that runs faster when you service it scores 1.
- Transfer. Can a competent outsider who has never seen your business predict the answer? Yes scores 1. If only insiders can, there is no expertise available to buy.1
- Reversal. Commit fully, be wrong, walk back through the door. Amazon’s split is the usable one: two-way doors are changeable and belong to small groups deciding fast, one-way doors are near-irreversible and earn slow deliberation.6 A two-way door scores 1.
Four or five: complicated. Buy the answer. Hire the specialist, copy the standard, run the outside-view forecast, sequence the work.
Zero to two: complex. Buy information instead. Cut the commitment into probes small enough to lose, run them in parallel, and decide what result would make you scale before launching any of them.
Three: treat it as complex. A probe wasted on a complicated problem costs you a week. A plan spent on a complex one costs you the quarter.
GEER: what to pull, and in what order
Misclassification does its damage through four channels: commitment size, how long the answer takes, whether you can walk it back, and attribution, meaning whether you can tell what caused the result. Start at the top.
- Instrument before deciding. Put a counter on the thing you are arguing about: orders per agent per week, repeat rate by source, float turnover. Hits attribution. Costs an afternoon.
- Run the repeat read. Same input, twice, a week apart, written down. Hits every channel, costs only patience.
- Size to affordable loss. Cut the commitment to what you could lose twice without changing the runway conversation.
- Write the amplify and dampen thresholds first. The number that makes you double, the number that makes you stop, both recorded before launch. Removes the argument you would otherwise have with yourself later.
- Run probes in parallel. Three at a third of the budget beats one at full size, because complex systems do not queue politely while you learn.
- Change what you pay for. Contracts that motivate exploitation look like standard pay-for-performance. Contracts that motivate exploration show substantial tolerance, or even reward, for early failure and pay for long-term success.9 Pay on percent of plan delivered and you will get plans.3 Costs a difficult conversation.
- Buy the complicated slice. Payroll compliance, tax filing, the payment integration. Hire those out at a fixed price and spend your attention on the part that adapts.
No-lever flag. Complex problem, and you cannot fund even one probe you can afford to lose. No technique closes that gap. You are in a financing-required state. Put that in writing and rebuild the quarter from there.
RADAR: the portfolio, dated
DO NOW, by T+3 days. Nothing here can hurt you, whichever way the reads land.
- Score the five reads, in writing, on your three largest open problems. Circulate them. Disagreement about a score beats the score.
- Take the lowest-scoring problem, cut its planned commitment to one third.
- Split that third into two probes that can run at once without contaminating each other.
- Write an amplify threshold and a dampen threshold for each probe. One number each. Send them to someone who will hold you to them.
HEDGE, by T+14. Small premiums, paid in case your scoring was wrong.
- Keep one specialist on a small retainer for the slices that scored complicated. Buying the answer stays correct wherever an answer exists.
- Publish one number per probe per week. Same number, same day, no commentary.
- Run the outside-view check: three comparable efforts outside your company, and what actually happened to them.10
DEFER AND TRIGGER. One-way doors. Hold them shut, and name today the signal that opens them.
- Defer: hiring the specialist team, signing the exclusive distributor, building the deep integration, announcing the pivot.
- Trigger to commit: one probe clears its amplify threshold in two consecutive weekly reports. Move the deferred spend behind it within seven days.
- Counter-trigger: by T+28 no probe has cleared and no read has changed score. That is a scoping failure, not a probe failure. Rewrite the problem statement before spending further.
Reading somebody else’s plan. Investors and boards get one question that does the work of ten: which reads did you score, and where is the amplify threshold written down. A Gantt chart with no threshold is a complicated-problem answer, and you can check whether it was earned.
CHAIN: what usually happens next
Look for precedent by mechanism, not by industry. Your match is any case where a designer pinned one variable inside an adaptive market and expected the rest to stay put: price caps, quotas, referral bounties, forced bundling. The CBK review notes that international experience with rate caps has in most cases produced undesirable outcomes, including reduced intermediation and transparency, reduced bank competition, and increased risk to financial stability.11 That is the base rate, and it is unkind.
Two present-state modifiers pull opposite ways in African markets. Cheap instrumentation shortens feedback and moves problems toward the readable end. High concentration lets one distributor or regulator override every local rule at once and move them back.
Second order: whoever loses from your rule builds a workaround, and the workaround becomes the baseline you now have to model. Third order: your team learns the metric is the target and starts producing the metric. Both showed up in Kenya, where banks kept the capped rate and changed the risk they would accept.11
Two corrections before this goes into a memo. Part of the Kenyan slowdown belongs to political risk ahead of the general election rather than to the cap, and the CBK cautions that outcomes visible in a single year may present a partial picture.11 A probe that succeeds while the whole market expands has told you less than it appears to. Check what the untreated part of your business did over the same window before crediting it.
Matrix-break flag. The classification is not permanent. Mobile money settlement data and instrumented agent networks give legible cues to environments that had none, so problems that were genuinely complex five years ago turn high-validity for whoever instruments them first. Re-score the reads every two quarters. The advantage belongs to whoever notices the flip early enough to stop probing and buy the answer.
What this ensemble cannot see
Four limits, all load-bearing.
The reads score your visibility, not the system. A problem looks complicated when you have not looked closely. Every score is provisional on your instrumentation, which is why lever one comes first.
Where the seam runs. Real problems are mixed. Distribution in Kampala is complex, the URA filing attached to it is complicated. Nothing here locates the join. You find it by cutting and watching which half stays still.
Probes nobody ran. Every result quoted above came from an experiment that got funded. Ideas killed in a meeting for being awkward leave no data, and they are the larger population.
Response time. No model here tells you how long a complex system takes to answer. A probe read too early looks exactly like a probe that failed, and killing a slow winner is the most expensive mistake in the method.
One action survives all four. Score the five reads today, on paper, and cut your single largest planned commitment to one third of its size by T+3. If the problem proves complicated you have delayed the answer by a week. If it proves complex you have bought back the quarter.
Sources and notes
- Kahneman, D., and Klein, G. “Conditions for Intuitive Expertise: A Failure to Disagree.” American Psychologist 64(6), 2009, 515-526. High-validity definition, the prolonged-practice and rapid-unequivocal-feedback conditions, Hogarth’s wicked environments and the fractionation of skill are all in the article’s summary and discussion sections. Full text.
- Schelling, T. C. “Dynamic Models of Segregation.” Journal of Mathematical Sociology 1, 1971, 143-186. The quoted conclusion, that systemic effects are overwhelming, that there is no simple correspondence of individual incentive to collective results, and that inferences about individual motives usually cannot be drawn from aggregate patterns, is in the abstract. Full text.
- Kohavi, R., Deng, A., Frasca, B., Walker, T., Xu, Y., and Pohlmann, N. “Online Controlled Experiments at Large Scale.” Proceedings of KDD 2013. The one-third figure for Microsoft, the Quicken Loans 33% quotation and the “percent of plan delivered” passage are in the Tenets section. Conference PDF.
- Camuffo, A., Gambardella, A., Messinese, D., Novelli, E., Paolucci, E., and Spina, C. “A scientific approach to entrepreneurial decision-making: Large-scale replication and extension.” Strategic Management Journal 45(6), 2024, 1209-1237. Four randomised control trials, 759 firms, 11,463 data points. Open-access full text, City St George’s, University of London.
- Mauboussin, M. J. “Untangling Skill and Luck.” Legg Mason Capital Management, 15 July 2010. The continuum, the sample-size rule and the reversion-to-the-mean rate are in the opening sections. Full text.
- Amazon.com, 2015 Letter to Shareholders. Type 1 and Type 2 decisions, one-way and two-way doors, and the footnote on organisations that apply the wrong process. Letter PDF.
- Hughes, N., and Lonie, S. “M-PESA: Mobile Money for the ‘Unbanked’.” Innovations: Technology, Governance, Globalization 2(1-2), 2007, 63-81. Pilot start date and eight agent stores, the list of unplanned uses observed during the pilot, the decision to commission researchers, and the March 2007 national launch are all in the authors’ own account. Article PDF.
- GSMA, State of the Industry Report on Mobile Money 2025. More than one billion registered accounts in Sub-Saharan Africa in 2024, twice the 2020 figure, and 75% above the level estimated in 2019 forecasts. Report PDF.
- Manso, G. “Motivating Innovation.” Author’s working paper dated 26 November 2010, marked on its first page as forthcoming in the Journal of Finance, where it appeared in 2011. The optimal contract for exploration exhibits substantial tolerance, or even reward, for early failure and rewards long-term success, in contrast to the pay-for-performance contract that motivates exploitation. Author PDF, Berkeley Haas.
- Flyvbjerg, B. “From Nobel Prize to Project Management: Getting Risks Right.” Project Management Journal 37(3), 2006, 5-15. Reference class forecasting, the outside view, and the American Planning Association endorsement. arXiv full text.
- Central Bank of Kenya, “The Impact of Interest Rate Capping on the Kenyan Economy,” March 2018. Law operational 14 September 2016; loan accounts down 26.1% between October 2016 and June 2017, large banks down 27.7%; MSME rationing estimated to have lowered 2017 growth by 0.4 percentage points; the international-experience summary and the partial-picture caveat are in the abstract and the report’s discussion. Report PDF.