A new distributor wants three months of stock on terms. You could give them a tenth of that and see what happens, and it feels weak, so you either give them the full order or you insist on cash up front and lose them.
Both of those throw away the only thing the first transaction is good for. The first order is not a small version of the deal. It is the instrument.
Why these three models
The decision is how large to make a first transaction with a counterparty whose type you cannot observe. The features that fire are private information about how patient they are, a signal about their conduct that is noisy rather than clean, and a test you get to design.
Three lenses, three error structures. Escalation produces an equilibrium answer about what structure separates the types. Noisy monitoring produces a complex answer about why you cannot simply watch and find out. The probe produces a random answer about how to design the trial so it carries information. The second is the reason the first is necessary, and the third is how you execute it.
1. Why small, and why escalating
The formal result is more specific than the folk advice about starting small, and the specificity is what makes it usable.
Set up the problem properly. Counterparties come in two types, patient and impatient, and the type is private. An impatient counterparty finds the immediate gain from defecting too tempting relative to the future, so even facing a partner who will never deal with them again, their best reply is to take what is available now. Where the chance of facing an impatient type is high enough, the only outcome is that nobody cooperates at all, and the good counterparties are shut out along with the bad.1
What breaks that deadlock is changing the scale over time. In the separating structure, patient players run at a deliberately reduced level for a stated number of periods, then move to full effort as long as their partner does too, abandoning the relationship after any shirking. Early periods are confidence building. Later periods are where the value is.1
The consequence founders resist is that the first period is not supposed to be profitable. It is a filter. Judging a pilot on its margin is judging a diagnostic on whether it made money, and it will always fail that test.
The second consequence is the one in the title. The structure works because the schedule is known. A counterparty who cannot see that the stakes rise on a stated timetable has no reason to sit through the small period, so the escalation has to be declared rather than merely intended.
Starting small, the separation lens
- Assumes: counterparties differ in patience, the difference is private, and the scale of the relationship can change over time.
- Fits because: you cannot tell the two types apart at the point of first contact.
- Breaks when: the counterparty has an outside option good enough that sitting through a small period is not worth it, in which case you screen out the strongest partners first.
- Evidence: grade A. Formally established, and matched by field evidence that relationship value rises with relationship age.
- Counteracts: judging a pilot on its margin.
- May reinforce: chronic under-committing, and a book of relationships that never escalate.
2. Why you cannot just watch and find out
The second lens explains why the structure is needed at all, rather than simple attentiveness.
In the clean version of a repeated relationship, you observe what your counterparty did. In the real version you observe a noisy signal of it. A delayed payment caused by their customer defaulting and a delayed payment caused by their choosing to hold your money arrive at you identically.
The consequence is structural rather than a matter of vigilance. Under imperfect monitoring, punishment phases occur along the equilibrium path, in periods when both sides know perfectly well that nobody deviated.1 You will penalise counterparties who did nothing wrong, and they will penalise you, and neither of you is being unreasonable.
This is exactly why the small first period is worth its cost. It is not that small transactions reveal character faster. It is that a small transaction bounds what a false positive costs you, so you can afford to act on a noisy signal at all. At full scale the same noise is unaffordable, so you hesitate, and hesitation is how a manageable dispute becomes an unmanageable one.
Noisy monitoring, the signal lens
- Assumes: you observe a noisy public signal of their conduct rather than the conduct.
- Fits because: the failures you will see are ambiguous between bad luck and bad faith.
- Breaks when: monitoring is private and each side sees something different, where coordination fails in ways this model does not describe.
- Evidence: grade A. Structural, and confirmed wherever outcomes are observable and effort is not.
- Counteracts: the belief that paying closer attention will resolve the ambiguity.
- May reinforce: excusing genuine bad conduct as noise.
3. Designing the trial so it tells you something
The third lens turns the structure into an experiment, and most pilots fail this test.
You cannot see inside your counterparty. So do not model their mechanism. Manipulate the input and read the output. Beer’s illustration is a child who has no idea how the side of his cot is constructed and who therefore adopts a black box strategy: he manipulates the inputs, discovers that shaking and rocking produces the collapse he wants, and gets the result without understanding anything.2
The half that founders skip is the design rule. The most efficient searching procedure is the one offering the highest entropy at each selection.2 A trial that could only come out one way carries no information. A first order small enough that anyone would honour it teaches you nothing. That is the standard pilot: sized to be safe, therefore sized to be uninformative.
So size the first period at the level where an impatient counterparty would find defection tempting and a patient one would not. That is the point of maximum information, and it is deliberately uncomfortable. Below it you learn nothing. Above it you are no longer running a test, you are exposed.
Black box probe, the design lens
- Assumes: you can vary an input and observe an output, and nothing about the internals.
- Fits because: their type is private and will not be revealed by asking.
- Breaks when: the response arrives outside your observation window, at which point this becomes superstition.
- Evidence: grade A. An experimental design policy resting on information theory rather than an empirical claim.
- Counteracts: pilots sized for comfort.
- May reinforce: treating a counterparty as a specimen, which they will notice.
The levers, cheapest first
- Publish the escalation schedule. Two orders at this level, then double, then standard terms. Written into the agreement. Costs nothing and it is what makes the small period tolerable to a good counterparty.
- Size the first period at the informative level, not the safe one. Ask what exposure would tempt an impatient counterparty. That is the number.
- Say the reason out loud. “We start every new distributor here and move on this timetable” reads as a policy. Silence reads as distrust of them specifically.
- Write down what would count as a failure before you start. Under noisy monitoring you will see something ambiguous, and the threshold decided afterwards is decided emotionally.
- Give them a way to accelerate. A counterparty who wants to post a deposit or provide a reference to skip a step is telling you their type, which is the whole object of the exercise.
What to do before the next new counterparty
Do now, sized at one hour, effect immediate. Write the standard escalation schedule for new counterparties: three steps, the trigger to move between them, and the exposure at each. Reversible, free, and dominant across every scenario about who you are about to deal with.
Hedge, where the premium is the whole loss, live before the next agreement. Run the next new relationship on that schedule rather than negotiating it fresh. If they were always going to be excellent you have delayed full volume by two cycles, and that is the entire downside.
Defer and trigger, size fixed now. Do not re-paper your existing relationships. Pre-commit the trigger: the first time an existing counterparty misses in a way you cannot explain, they move back one step on the schedule automatically, without a meeting. Fix what one step means now, because a de-escalation designed during a dispute is designed under pressure.
Watch the arrivals. The schedule lands this week. What it tells you about a counterparty lands only after enough periods for the escalation to bite, which is months, and the temptation to skip ahead will peak in month two.
What usually happens next
A counterparty who asks to skip a step is worth listening to closely. Under noisy monitoring you cannot read their type from conduct alone, so a costly voluntary signal, a deposit, a personal guarantee, an audited reference, is one of the few clean pieces of information available. It is also the moment where an impatient counterparty reveals itself by refusing.
Run the break test first. Has a rule changed, has an actor entered or left, has a measurement become a target? If a competitor has entered offering full terms immediately, your small first period is now priced against a real alternative, and the structure that separated types last year screens out good counterparties this year.
If nothing broke, the base rate is encouraging and specific. Relationship value rises with relationship age, and reliability observed at the moment of a shock predicts both future survival and future relationship value.3 The escalation is not merely a defence. It is building the thing that makes the relationship worth having.
There is a second pattern in how these arrangements fail, and it is not the one founders fear. The common failure is not a counterparty who defects in the small period. It is a small period that never ends, because nobody owned the escalation and the relationship simply settled at step one. The schedule needs a date and a name attached, or it becomes a permanent ceiling that both sides quietly accept.
Subtract the counterfactual before crediting your screening. A counterparty who performed well through a small first period in a quiet quarter has not been tested. The information came from a period where defection was tempting, and if none was, you learned nothing and should not escalate on it.
What this ensemble cannot see
All three lenses treat you as the one running the test. Your counterparty is running one too.
A small first order tells them something about you: that you are cautious, that you have process, and possibly that you are small enough to need the protection. Sophisticated counterparties read escalation schedules as information about the company proposing them, and nothing here models what they conclude or how it changes their behaviour toward you.
There is also a genuine cost this framework understates. Screening on patience screens out counterparties with good outside options, because sitting through a deliberately unprofitable period is most expensive for the ones who have alternatives. The structure is best at separating types among counterparties who need you, which is a biased sample of the market.
And one property none of these models contains: the escalation schedule becomes a norm, and norms are read as ceilings. Counterparties who have been through three steps often stop asking for more, and you will find yourself with a book of relationships that all sit at step three because nobody remembered there was a step four.
The one action that survives the ignorance: before the next new counterparty, write the exposure at which an impatient party would be tempted and a patient one would not. If you cannot name that number, your first order is sized for comfort and will teach you nothing, whatever it does for your risk.
Who has to move
This only changes anything if whoever negotiates new terms reads it, and they are usually measured on volume signed rather than on counterparties correctly sorted. The cheapest first test is the written three-step schedule applied to one new relationship. It costs an hour, and it converts an argument about trust into a policy, which is far easier to hold.
Sources and notes
- George J. Mailath and Larry Samuelson, Repeated Games and Reputations: Long-Run Relationships, Oxford University Press, 2006. Starting small is section 5.2.4, following Watson, including the setup in which impatient types shirk immediately even against a partner playing grim trigger, the result that a sufficiently high probability of the impatient type produces perpetual shirking, and the separating structure in which patient players choose moderate effort for a stated number of periods before escalating, with early periods described as confidence building. Imperfect public monitoring and the finding that punishments occur along the equilibrium path in periods when both sides know nobody deviated are section 7.2.1. The effect of increased monitoring precision on the achievable set is section 7.4.
- Stafford Beer, Cybernetics and Management, English Universities Press, 1959. The black box strategy, the child and the cot, and the argument that the most efficient searching procedure is the one offering the highest entropy at each selection are in the chapter on the Black Box. Note: the copy consulted is an image-only scan read via optical character recognition, so it is cited qualitatively and no figure is quoted from it.
- Rocco Macchiavello and Ameet Morjaria, The Value of Relationships: Evidence from a Supply Shock to Kenyan Rose Exports, University of Warwick Economic Research Paper 1032, published in the American Economic Review. https://wrap.warwick.ac.uk/id/eprint/59370/1/WRAP_twerp_1032_macchiavello.pdf. The abstract states that the value of the relationship increases with the age of the relationship, that during an exogenous negative supply shock sellers prioritise relationships consistently with the model, and that reliability at the time of the shock positively correlates with future survival and relationship value.
A note on the uncomfortable part. The informative first order is larger than the safe one, by construction, because a test nobody could fail carries no information. That is a real tension with the survival constraint, and where the two conflict the survival constraint wins. A probe sized above what you can afford to lose is not a probe, it is the exposure you were trying to avoid.
Joshua Agonya Pi’Rwot, Founder.