Your team keeps over-ordering, then under-ordering. Your hiring runs hot, then freezes. Your cash forecast is wrong in the same direction three quarters running. So you do the three things anyone would do: you hire more senior people, you put a bonus on the number, and you tell yourself the market will discipline whoever is getting it wrong.
All three have been tested against this exact failure. None of them moved it.
Why these three models
The decision is whether to spend your next intervention on people or on structure. The features that fire are a quantity that accumulates, a delay between action and response, a problem that keeps returning, and a process that is failing while everyone works hard.
Three lenses with genuinely different error structures. The first says the operator will misread the system and produces a complex, non-equilibrium answer. The second asks whether the process can be controlled at all and produces an equilibrium answer. The third refuses your particulars and asks what usually fixes this class of failure, which produces a cycle-regime answer. All three are grade A, which is unusual, and it is why this piece can be blunt.
1. The constant: people misread feedback, and it does not wear off
Start with the finding, because it is stronger than almost anything in behavioural economics and it is much less quoted.
People adopt an event-based, open-loop view of causality. They ignore feedback processes, fail to appreciate the delays between an action and its response, do not understand stocks and flows, and are insensitive to the nonlinearities that change which loop dominates as a system evolves.1 That is a list of ordinary cognitive limits, and on its own it would be unremarkable.
What makes it remarkable is the next sentence in the source. These misperceptions are robust to experience, to financial incentives, and to the presence of market institutions.1
Read that as the list it is. It is exactly the list a sceptic would produce when told that a laboratory result will not survive contact with real operators. Experience will fix it. Money will fix it. Competition will fix it. Each was tested. Each failed.
The supporting record is specific rather than atmospheric. In the Beer Distribution Game, a four-node production and distribution chain, players generate costly oscillation with average costs more than ten times greater than optimal, and the range of players runs from high school students to chief executives.1 Customer demand in that game is constant. The cycles are manufactured entirely inside the participants’ own decisions. In a separate experiment, subjects responsible for capital investment in a simple model of an economy generate large-amplitude cycles while consumer demand is held flat.1
And it gets worse rather than better as the problem gets harder: the greater the dynamic complexity of the environment, the worse people do relative to what was achievable.1
A separate line of experiments isolates the simplest possible version and finds the same thing. Asked to infer the behaviour of a plain stock and flow system, with no strategy and no counterparty, highly educated adults fail, and the failure is not attributable to an inability to interpret graphs, to contextual knowledge, to motivation, or to cognitive capacity.3 The authors call it stock-flow failure and treat it as a fundamental reasoning error rather than an artefact of how the task was presented.3
Put the two together and the excuse space closes. It is not the incentives, because incentives were tested. It is not the seniority, because chief executives were tested. It is not the presentation, because presentation was tested. It is the structure, and structure is the only thing left to change.
Feedback misperception, the operator lens
- Assumes: the decision maker’s causal map is far simpler than the system, and they cannot infer the dynamics of even simple maps.
- Fits because: the failure recurs under different people, which rules out the people.
- Breaks when: the task is detail-complex rather than dynamically complex. On a large combination space with fast feedback, expertise and incentives work fine.
- Evidence: grade A. Replicated across many experiments and robust to experience, incentives and market institutions.
- Counteracts: the instinct to solve an operating failure by changing who is doing it.
- May reinforce: fatalism, and an excuse for genuinely poor execution.
Notice what this model is for. Most models tell you what will work. This one tells you what will not, on strong evidence, and that is rarer and more valuable. It removes three of the four options you were considering.
2. The test: can this process be controlled at all
The second lens asks a question most operating reviews never ask.
Count the distinguishable states the thing you are managing can be in. Then count the distinguishable responses your process actually has. If the second number is smaller than the first, the process cannot regulate the system, however hard anyone works and however many people you add.
This is the law of requisite variety, and Beer adopts it as an axiom of controlling complex systems: only where there is enough permutative variety to give a one-to-one transformation from the controlled system into the control box does requisite variety exist.2
There are exactly two repairs and no third. Attenuate the variety arriving at you: standardise the offer, template the intake, restrict what a customer can request without escalation. Amplify your own: add response categories, delegate authority so more decisions can be made without routing upward, automate a class of cases so the humans handle the rest.
The reason this belongs next to the first model is the arithmetic. Adding headcount with an unchanged response set changes neither count. Your team of four with three canned answers becomes a team of eight with three canned answers. The states have not fallen and the responses have not risen, so the queue does not clear, and everyone reasonably concludes that they are still understaffed.
Requisite variety, the controllability lens
- Assumes: regulation requires the controller to have at least as many distinguishable responses as the controlled system has states.
- Fits because: the process is failing while effort is high, which is the signature.
- Breaks when: nothing. It is close to analytic, which is its own limitation.
- Evidence: grade A, and honestly so: it is near-definitional, so it cannot be false and it will never surprise you.
- Counteracts: treating every capacity problem as a staffing problem.
- May reinforce: false precision, because counting distinguishable states is a judgement not a measurement.
3. The reference class: what actually fixes this
Before estimating what will work in your case, ask what works in this class of case.
The pattern in the record is consistent and unflattering. Interventions aimed at the person, meaning hiring, training and incentive design, have been tested directly against dynamic-complexity failures and have not moved them. Interventions aimed at the structure have. Shortening the delay between an action and its observable result changes behaviour. Making a stock visible where it was previously inferred changes behaviour. Writing down an explicit ordering rule, so the decision stops being a judgement made under pressure, changes behaviour.
The base rate here is not a percentage. It is a direction, and the direction is that the people-shaped fix has a poor record against this specific failure mode and an excellent record against others. Which is the trap: the fix that works everywhere else is the one you will reach for first, and it is the one with the evidence against it here.
Reference class, the history lens
- Assumes: failures matched on structure respond to the same class of intervention.
- Fits because: this failure has occurred in thousands of operating teams.
- Breaks when: the rules changed, in which case the historical record describes a different process.
- Evidence: grade A as a method, and the specific base rate here is directional rather than numeric.
- Counteracts: the inside view, which builds the answer from your team’s particular personalities.
- May reinforce: importing a base rate from a firm with different information systems.
Run the structure before you run the people
Take the failing process and answer five questions in writing. They take an hour.
- Position the boundary. What is inside this analysis and what is outside? Write down what you left out, because that list is what you have agreed to be surprised by.
- Units. Is the thing you are trying to move a stock or a flow? You cannot set a stock. If you are trying to move one, name the flow you are actually touching.
- Lag. How long between the action and the observable response, and is your correction sized as though the lag were zero? That combination is the standard oscillator and it will produce the swing you are already seeing.
- Structure. Which loop dominates, the reinforcing one or the balancing one, and what shape does that produce over time?
- Excluded. What did the boundary leave out? Those are the effects you did not plan for. They are not side effects. There are no side effects, there are just effects, and the unexpected ones are a report on the boundary you drew.
What to change this month
Do now, capped at one working day, effect visible within two weeks. Take the failing process and count two numbers: how many distinct situations arrive, and how many distinct responses exist. If responses are fewer, you have your answer without any further analysis. Reversible, cheap, and dominant across every scenario about whose fault it is.
Hedge, where the premium is the whole loss, cover live before the next cycle. Shorten one reporting delay by half for a single period. If the oscillation was not delay-driven you have spent a little reporting effort, and that is the entire downside.
Defer and trigger, size fixed now. Do not restructure the whole operating model this quarter. Pre-commit the trigger instead: the next time the same correction has to be made twice, an explicit ordering rule gets written for that decision, and it gets written before anyone is hired. Decide the scope of that rule now, while nothing is on fire.
Note the dates on those, because they are the point. A DO NOW whose effect arrives after the DEFER trigger fires is misordered, and that misordering is how a sensible plan produces the oscillation it was meant to fix.
What usually happens next
Run the break test before the base rate. Has a rule changed, has an actor entered or left, has a measurement become a target? If the number you are about to stabilise is one the team is now compensated on, the historical relationship between that number and performance is describing a different process, and the base rate does not transfer.
If nothing broke, name the shape before explaining it. Oscillation with roughly constant amplitude points at a delay and a gain. Growth that flattens points at a balancing loop arriving. Growth that flattens and then falls points at overshoot, and overshoot is the one worth catching early because its correction is not gentle.
Subtract the counterfactual before you credit the new hire. A process that stabilised after a senior appointment may have stabilised because the appointment coincided with a quieter quarter. The test is whether the underlying delay changed, and usually it did not.
What this ensemble cannot see
These three models can tell you that a structure will oscillate. None of them can tell you the amplitude or the date, and the honest reason is in this same literature: unaided mental simulation of feedback structures fails even for experts, and the formal models built to do it are not predictive in the statistical sense. They reproduce modes of behaviour rather than forecasting values. So take the shape from this analysis and never take a number.
There is a second limit, and it cuts against the whole argument. This applies to dynamically complex failures: feedback, delays, accumulation. On a detail-complex problem, meaning a large space of combinations with fast honest feedback, experience and incentives work exactly as you would expect. Applying this article’s conclusion to a task like scheduling or pricing lookups would be a straightforward misuse, and you would be firing a good intervention for the wrong reason.
And one property none of these models contains: as you standardise to raise variety, you will lose the informal handling that was quietly absorbing the awkward cases. That absorption does not appear in any stock, flow or response count, and you will notice it only when it stops.
The one action that survives the ignorance: before your next operating review, take the last three problems that had to be fixed twice, and for each one write down the delay between the action and the visible result. If two of the three have a delay longer than your review cycle, the review cycle is the intervention, and no hire will substitute for it.
Who has to move
The person who needs this is usually the one who has just been given budget to fix the problem, and budget arrives shaped like headcount. The cheapest first test costs an hour and no money: count the states, count the responses, and bring both numbers to the meeting where the hire is being discussed. If responses are fewer than states, the hire will not work and now you can say why.
Sources and notes
- John D. Sterman, Business Dynamics: Systems Thinking and Modeling for a Complex World, McGraw-Hill. Chapter 1, “Learning in and about Complex Systems”. The list of misperceptions and the statement that they “are robust to experience, financial incentives, experience, and the presence of market institutions” appear in section 1.3 alongside Table 1-4, which summarises the experimental studies. The Beer Distribution Game result, that players from high school students to CEOs generate costly fluctuations with average costs more than ten times greater than optimal, and the multiplier-accelerator result with constant consumer demand, are both listed there. The distinction between dynamic and detail complexity is in section 1.3.1. “There are no side effects, there are just effects” is in section 1.1.2, on policy resistance. The modelling process, with problem articulation and boundary selection as step one, is chapter 3.
- Stafford Beer, Cybernetics and Management, English Universities Press, 1959. The law of requisite variety, attributed to Ashby, is adopted in the chapter on the Black Box as an axiom of controlling complex systems, with the condition stated as sufficient permutative variety to provide a one-to-one transformation from the controlled system into the control box. Note: the copy consulted is an image-only scan and was read via optical character recognition, so it is cited qualitatively here and no figure is quoted from it.
- Matthew A. Cronin, Cleotilde Gonzalez and John D. Sterman, Why don’t well-educated adults understand accumulation? A challenge to researchers, educators, and citizens, Organizational Behavior and Human Decision Processes 108(1), 2009, pages 116 to 130. Author copy: https://www.mit.edu/~jsterman/CroninGonzalezSterman061210.pdf. The abstract states that highly educated people are often unable to infer the behaviour of simple stock-flow systems, that poor performance has been ascribed to artefacts including complex information displays, lack of contextual knowledge, the cognitive burden of calculation or the inability to interpret graphs, and that in a series of experiments the persistent poor performance is not attributable to an inability to interpret graphs, contextual knowledge, motivation, or cognitive capacity.
A note on what is deliberately absent. This piece does not tell you how large the oscillation will be or when it will peak. Models of this kind reproduce behaviour modes and do not forecast values, and a number lifted out of one and dropped into a plan would be a misuse of the source. Take the shape.
Joshua Agonya Pi’Rwot, Founder.