Orders from a third of your distributors dropped last month. The team has four theories. Two people have strong opinions about the algorithm, one blames pricing, one blames a competitor. The meeting has now happened twice.
None of you can see inside the thing you are arguing about. Stop trying to understand it and start manipulating it.
Why these three models
The decision is whether to keep investigating or to run a test, and if a test, how to design it. The features that fire are a system whose internals are genuinely inaccessible, a response that arrives well after the action, and a real cost to running the experiment at all.
Three lenses. The probe produces a random answer about how to act without understanding. Delay produces a complex answer about the lag that quietly ruins most tests. The value of information produces an equilibrium answer about whether the test earns its cost. The first two pull against each other usefully: one says act now and read the output, the other says the output will not arrive for a while, so design for that.
1. Do not model the mechanism
Beer’s treatment of the black box is the oldest good advice on this and the most concrete.
Where a system is inaccessible, you do not need to understand it in order to control it. His illustration is a small child who has no idea how the side of his cot is constructed. The system is inaccessible to him because of his own intellectual limitations, so he adopts a black box strategy: he manipulates the inputs, and he discovers that by a combination of shaking and rocking he can obtain the desired output, which is the collapse of the side of the cot.1
Beer’s point is that this is not a childish method. It is the method, and he applies it to exactly the systems founders face: the economy, the industrial company. For these, he writes, the box is absolutely black, and the techniques we should use to handle exceedingly complex systems are those appropriate to a black box.1
He also aims the same argument at the accounting most companies use to explain their own results, and it lands harder than anything written since. Of conventional cost analysis he says that from the point of view of cybernetic control the whole approach is basically wrong, because it deals with a homomorphism of the real situation, one in which cause-effect relationships are assumed to hold. But they do not hold.1
Attributing your order decline to a line in a spreadsheet is that error precisely.
Black box probe, the inaccessibility lens
- Assumes: you can vary an input and observe an output, and nothing about the internals.
- Fits because: the thing you are arguing about cannot be inspected by anyone in the room.
- Breaks when: the response arrives outside your observation window, at which point this degrades into superstition. That is model two.
- Evidence: grade A. An experimental design policy resting on information theory rather than an empirical claim about behaviour.
- Counteracts: theorising about a mechanism you will never see.
- May reinforce: testing trivia, because a clean test on an unimportant variable still feels like rigour.
2. Design the split so it can come out either way
The half that founders skip is not whether to test but how to choose what to vary, and Beer gives the rule.
The most efficient searching procedure is the one offering the highest entropy at each selection.1 He works the arithmetic: if you must find one item among ten and you examine each in turn, the entropy of that selection is low, and to raise it toward its maximum you have to change how you split the space rather than how hard you look.1
Translated: a test whose result you can predict carries no information, however carefully it is run. Sending your revised terms to the next twelve distributors is not a test, because there is no outcome that would change your mind. Sending version A to six and version B to six is a test, and it is a better one the more nearly it splits your uncertainty in half.
That is also the reason to test the thing you are least sure about rather than the thing you most want to be true. The comfortable variable is the one where you can guess the answer, which is exactly the one worth nothing.
3. The lag that ruins it
The second model is why most well-designed probes still fail, and it has nothing to do with the design.
A distributor’s reorder decision does not respond to your terms this week. It responds after their current stock runs down, after their own customers react, after their finance person notices. The response arrives weeks or months later, and by then you have changed two other things.
Correcting against a delayed signal at full strength produces overshoot, and this is where the two models collide productively. The probe says vary the input and read the output. The delay says you will read an output that reflects an input from two changes ago unless you hold still.
So the operational rule is unglamorous, and it follows directly from how delayed systems behave.3 Change one thing. Then wait longer than feels reasonable, on a date you wrote down before you started, and change nothing else in between. Most companies cannot do the third part, which is why most of their tests are uninterpretable rather than wrong.
Expect to underrate this specifically. Reasoning about delay and accumulation defeats highly educated adults in controlled experiments, and the failure is not attributable to graph literacy, contextual knowledge, motivation or cognitive capacity.2
Delay and oscillation, the lag lens
- Assumes: the response to your change arrives after a lag you did not choose.
- Fits because: the system responding is a chain of people and inventories, not a server.
- Breaks when: the response is genuinely immediate, where holding still is unnecessary caution.
- Evidence: grade A. Structural, and the human failure to reason about it is well replicated.
- Counteracts: reading an early result as the result.
- May reinforce: waiting on a test that has genuinely already failed.
4. Whether the test is worth running
The third lens stops the first two from turning into a testing culture that measures everything and decides nothing.
Information has value only where it can change an action. Before designing anything, write down what you would do under each result. If a positive result and a negative result lead to the same decision, the test has zero value regardless of how elegantly it splits the space, and the effort belongs elsewhere.
This also sizes the test honestly. If detecting the difference you care about would need more distributors than you have, the correct conclusion is that you cannot detect it. Saying so is a legitimate and cheap answer, and it is far better than an underpowered split whose noise you will then interpret as a finding.
A test you would overrule is not a test. It is a delay with a spreadsheet attached.
Value of information, the decision-relevance lens
- Assumes: a test is worth its cost only if some result would change what you do.
- Fits because: the test consumes real customers and real time.
- Breaks when: the value is political rather than decisional, and the test exists to settle an argument between people.
- Evidence: grade A. A formal result rather than an empirical claim.
- Counteracts: running experiments whose outcome you have already decided to ignore.
- May reinforce: refusing to measure anything whose answer might be uncomfortable.
The levers, cheapest first
- Write what you would do under each result, first. Two lines. If they match, cancel the test and save the customers.
- Pick the variable you cannot predict. Not the one you hope about. The one where you would genuinely bet either way.
- Split the population, do not sequence it. Before and after is contaminated by everything else that changed. A and B at the same time is not.
- Write the read date before you start. Long enough for the slowest part of the chain to respond, decided while you are calm.
- Change nothing else in the window. This is the hardest one and the one that actually decides whether you learn anything.
- Stop the meeting. A second meeting about why something happened, with no new data between them, is a meeting that cannot produce anything.
What to do this week
Do now, sized at twenty minutes, effect immediate. Take the thing your team has theorised about twice and write the two lines: what you would do if the answer is A, what you would do if it is B. Reversible, free, and dominant across every scenario about which theory is right.
Hedge, where the premium is the whole loss, live before the next cycle. Split one segment and vary one term. If the theory you favoured was right you have lost a little consistency across accounts for one period, and that is the entire downside.
Defer and trigger, size fixed now. Do not rebuild how the company experiments. Pre-commit the trigger: the next time a question survives two meetings without new data, it becomes a split test automatically rather than a third meeting. Decide now who is allowed to call that, because a rule with no enforcer is a preference.
Note the arrivals. The two lines land today. The split can start this week. The reading cannot happen until the lag has run, and the entire value of the exercise depends on not looking before then.
What usually happens next
Run the break test first. Has a rule changed, has an actor entered or left, has a measurement become a target? A competitor entering during your test window contaminates it completely, and the honest response is to void the test rather than to interpret it.
If nothing broke, the pattern is consistent. The test gets started, something urgent changes a second variable in week two, and the result is read anyway because the effort has been spent. That is worse than not testing, because it produces a confident wrong belief that then gets defended.
Subtract the counterfactual before crediting the change. If orders recovered in the arm you changed and also in the arm you did not, the recovery was seasonal and you have learned nothing about your terms, which is itself a real finding.
What this ensemble cannot see
All three lenses treat the counterparty as a system to be probed. They are people, and they notice.
Distributors talk to each other. A split test that gives six accounts better terms than six others is discoverable, and the discovery costs trust in a way no model here prices. In a market where your counterparties know each other, which describes most distribution networks, the probe is not invisible and running it has a cost that does not appear in the design.
There is also a limit on the central method. The black box approach tells you which input moves the output. It never tells you why, which means it cannot tell you whether the relationship will survive a change in conditions. You get a lever without a mechanism, and levers without mechanisms stop working without warning.
And one property none of these models contains: the discipline of changing one thing and waiting is organisationally almost impossible under pressure. The month you most need to run a clean test is the month you least can, and no amount of understanding the method changes that.
The one action that survives the ignorance: before the next meeting on this question, write the two lines about what you would do under each answer. If they are the same line, cancel the meeting. If they differ, the meeting should be about designing the split, not about which theory is right.
Who has to move
The person who needs this is whoever keeps calling the meeting, and the meeting feels productive because the theories are getting sharper. The cheapest first test costs twenty minutes: the two lines, written before anyone speaks. If they match, you have saved the meeting and everything after it.
Sources and notes
- Stafford Beer, Cybernetics and Management, English Universities Press, 1959. The black box chapter carries all of the following: the child and the cot, described as adopting a black box strategy by manipulating the inputs to obtain the desired output; the argument that for essentially scientific systems such as the economy or the industrial company the box is absolutely black, and that the methods appropriate to exceedingly complex systems are black box methods; the entropy argument that the most efficient searching procedure is the one offering the highest entropy at each selection, with the worked ten-item example; and the criticism of conventional cost analysis as dealing with a homomorphism of the real situation in which cause-effect relationships are assumed to hold but do not. Note: the copy consulted is an image-only scan and was read via optical character recognition, so it is cited qualitatively here and no figure is quoted from it.
- Matthew A. Cronin, Cleotilde Gonzalez and John D. Sterman, Why don’t well-educated adults understand accumulation? A challenge to researchers, educators, and citizens, Organizational Behavior and Human Decision Processes 108(1), 2009, pages 116 to 130. Author copy: https://www.mit.edu/~jsterman/CroninGonzalezSterman061210.pdf. The abstract states that highly educated people are often unable to infer the behaviour of simple stock-flow systems, and that persistent poor performance is not attributable to an inability to interpret graphs, contextual knowledge, motivation, or cognitive capacity.
- John D. Sterman, Business Dynamics: Systems Thinking and Modeling for a Complex World, McGraw-Hill. The treatment of delays between action and response as the standard source of oscillation, and the consequence that correcting at full strength against a delayed signal produces overshoot, is developed across chapters 4 and 17. Used in section 3 for the requirement to fix a read date in advance and hold every other variable still until it arrives.
A note on citing a 1959 book about algorithms. Beer was writing about industrial plants and national economies, not recommendation systems, and the specific technologies he had in mind are obsolete. The argument transfers because it never depended on the technology: it depended on the system being inaccessible to the person trying to control it, which is more true of a modern platform than of anything he was looking at.
Joshua Agonya Pi’Rwot, Founder.