You built a term sheet response designed to signal confidence without appearing to signal confidence. The investor read it as a slow reply. Not because they are unsophisticated, but because almost nobody runs the regress, and the measured answer is that most people stop at one or two steps.
The design rule that follows is uncomfortable and immediately useful. Build for a counterparty doing one step. Anything that requires them to reach the third step is a cost you pay and they never receive.
Why these three models
The decision is whether to make a move whose meaning depends on someone else’s inference. The features that fire are strategic actors with conflicting interests, information asymmetry, and an audience of readers who form judgements independently.
Three lenses with different error structures. Reasoning depth tells you how far the inference will actually travel, and it produces a complex, non-equilibrium answer. Signalling tells you what a costly action transmits regardless of depth, and it produces an equilibrium answer. The black box probe tells you how to act when you still cannot see inside the decision, and it produces a random answer about which test to run. Depth and signalling disagree in a productive way, which is the point of running both, and the probe is what you do when they leave you uncertain.
1. Depth: how far the inference travels
The level-k model starts from an anchor. A level-0 player does not reason strategically at all and either picks the salient option or picks at random. A level-1 player best responds to level-0. A level-2 player best responds to level-1. And so on, in principle without limit.
In practice it stops almost immediately. In the beauty contest, where players guess a target fraction of the average guess, the level-k prediction is a clean formula and the observed guesses cluster on the level-1 and level-2 predictions rather than on the equilibrium.1, 5 Across the matrix games surveyed in the same source, roughly ninety per cent of the population splits between level-1 and level-2.1
The cognitive hierarchy refinement makes it more realistic without changing the conclusion. Instead of assuming everyone else sits at exactly one level below you, a player believes there is a spread of lower types and holds correct beliefs about their relative frequencies, while remaining oblivious to anyone at or above their own level.1, 6 The formal version fits a single parameter, the average number of thinking steps in the population, and estimating it across many games is what turns “they reason less far than you think” into a number you can design against.6 Everyone believes they are one step ahead. That belief is what caps the depth.
Reasoning depth, the inference lens
- Assumes: players best respond to an assumed lower level rather than to an equilibrium.
- Fits because: your move only works if someone infers something from it.
- Breaks when: the setting is an auction. See the failure declared below, which is not a footnote for a founder.
- Evidence: grade B plus. Applied across more than a hundred games with process evidence from eye-tracking and response times, and a documented failure in auctions.
- Counteracts: the assumption that a rational counterparty will work it out.
- May reinforce: condescension, and an under-estimate of a professional operating in their own domain.
The failure has to be stated because it sits inside the fundraising process itself. In a controlled test built specifically to separate level-k from equilibrium in auctions, the level-k model substantially under-predicted observed bids and was clearly outperformed by equilibrium. Fitting it to the data produced implausibly high estimated levels which bore no relation to the levels the same subjects showed in a game known to trigger level-k reasoning. And subjects almost never appealed to iterated reasoning when asked to explain how they bid.2
So the honest position is narrow and worth holding precisely. Use depth to design a pitch, a price and a term. Do not lean on it inside a competitive bidding process, because that is where it has been tested and lost.
2. Signalling: what a costly move transmits anyway
The second model is the reason all is not lost. Signalling does not require the receiver to run an inference chain. It requires the signal to be costly in a way that a weaker type could not afford.
The cleanest demonstration is unravelling under verifiable disclosure. When restaurants in Los Angeles were given hygiene grade cards without a requirement to display them, the A-rated restaurants displayed theirs to avoid being taken for B. Once A restaurants displayed, B restaurants displayed to avoid being taken for C. Eventually even C restaurants displayed.1 Nobody in that chain performed four steps of reasoning. Each party performed one step against an observable fact.
That case is not a parable, and the measured outcome is stronger than the story. Studying the Los Angeles ordinance directly, Jin and Leslie found the grade cards caused restaurant health inspection scores to increase, consumer demand to become sensitive to changes in hygiene quality, and foodborne illness hospitalisations to decrease.4 A verifiable one-step signal changed what firms actually did, not merely what buyers believed.
That is the design principle. A signal that works at depth one is one where the absence of the signal is the information. You do not need your investor to reason about why you disclosed. You need the market state where not disclosing is what carries meaning.
And this is exactly where verifiability decides everything. Where disclosure cannot be verified, the unravelling never starts, no type separates, and the good operator is priced as the average one.1 The cost of that pooling is not a communication problem and no amount of narrative repairs it.
There is a further constraint that the two models produce jointly, and neither produces alone. Unravelling is a chain of inferences, so it runs exactly as many rounds as the audience is willing to run and then it stops. The same source shows this directly in film releases, where studios withhold a picture from critics before opening weekend and audiences only draw the inference so far, so the withholding survives above a quality level that a fully reasoning audience would have priced immediately.1 Read against the depth evidence, that gives a bound worth carrying: disclosure pressure in your market is not set by what logic implies, it is set by how many steps your buyers actually take. In a market of one-step readers, the second-worst operator is safe.
Which reframes what a verification layer is for. It is not there to reward the good operator. It is there to shorten the chain, so that a single observable fact does the work that four rounds of inference would otherwise have to do and never will.
Signalling, the transmission lens
- Assumes: a signal separates types only when a weaker type could not bear its cost.
- Fits because: you hold private information the other side would pay to have.
- Breaks when: disclosure is not verifiable, in which case all types pool and silence costs nothing.
- Evidence: grade A. Robust theory with clean field demonstrations.
- Counteracts: the belief that clearer explanation solves a credibility problem.
- May reinforce: expensive signalling in markets where nobody is reading.
3. The probe: stop modelling their mechanism
The first two lenses tell you that your counterparty reasons less far than you assumed, and that a costly, verifiable signal transmits anyway. The third tells you what to do when you still cannot work out what is going on inside their decision.
The answer is to stop trying. Where a system is inaccessible, do not model the mechanism. Manipulate the inputs and read the outputs. Beer’s illustration is a small child who has no idea how the side of his cot is constructed, and who therefore adopts a black box strategy: he manipulates the inputs, discovers that a combination of shaking and rocking produces the collapse he wants, and gets the result without ever understanding the mechanism.3
The second half of that idea is the part founders skip. The search procedure matters as much as the searching. “The most efficient searching procedure is the one offering the highest entropy at each selection.”3 Testing candidates one at a time is a low-information procedure: most trials come out the way you expected and you learn almost nothing. Splitting the space so that each trial could genuinely come out either way is the high-information one.
A test that can only come out one way tells you nothing. That is the whole standard, and most founder experimentation fails it. Sending a revised deck to the next twelve investors is not a test, because there is no version of the result that would change your mind. Sending version A to six and version B to six is a test.
Beer aims the same point at the accounting most companies use to explain their own outcomes. Conventional cost analysis, he writes, “deals with a homomorphism of the real situation, one in which cause-effect relationships are assumed to hold. But they do not hold.”3 Attributing a lost round to a line in your deck is that error exactly.
Black box probe, the inaccessibility lens
- Assumes: you can vary an input and observe an output, and nothing about the internals.
- Fits because: the first two lenses established that you cannot rely on reading their reasoning.
- Breaks when: the response arrives outside your observation window, at which point this degrades into superstition.
- Evidence: grade A. An experimental design policy resting on information theory, not an empirical claim about people.
- Counteracts: theorising about a decision process you will never see.
- May reinforce: testing trivia, because a well-split test on an unimportant variable still feels like rigour.
The levers, in order of what they cost you
- State the thing. Replace every implication with a sentence. If your deck implies retention is strong, write the retention number. Cost: an afternoon. Fully reversible, and it is the only lever that works at depth one.
- Make absence informative. Publish the metric that your weaker competitors cannot publish. The signal is not that you disclosed. The signal is that they did not.
- Buy verifiability where it does not exist. A third party who can confirm the number at low cost converts your disclosure from cheap talk into a separating signal. This is a purchase, not an argument.
- Split the next batch. Where you are about to send the same thing to twelve people, send two versions to six each and vary one element. A test that cannot come out against you is not a test.
- Stop paying for depth three. Audit your last three strategic moves for one that required the counterparty to infer your inference about their inference. Cut it. Nobody received it.
What to do before the next pitch
Do now, capped at an afternoon. Take your current deck and mark every claim that requires an inference rather than stating a fact. Convert the top five. This is reversible and it dominates across every scenario about who is reading.
Hedge, where the premium is the whole loss. Publish one number your competitors will not publish. If the market is not reading, you have given away a small amount of information advantage, and that is the entire downside.
Defer and trigger, with the size set now. Do not buy third-party verification today for every metric. Pre-commit the trigger: the first time a serious counterparty discounts a number you gave them, that specific number goes to a verifier. Decide now what you are willing to spend on that, so the decision is not made in the moment after a rejection.
Every one of those carries a cap, and the cap is set before the opportunity appears. A reversible move sized at a quarter of your remaining runway is not a reversible move.
What usually happens next
Run the break test before the base rate. Has a rule changed, has an actor entered or left, has a measurement become a target? If the metric you are about to disclose is one the market has started optimising for, then the historical relationship between that metric and quality is describing a different process, and disclosing it buys you less than it used to.
If nothing broke, the reference class is stable and unglamorous. Moves that require deep inference get read as noise. Moves that state a fact get read as a fact. Founders systematically over-estimate how much of their intent survives the trip, because they know what they meant and cannot un-know it.
Subtract the counterfactual before concluding your subtle move worked. If a round closed after you made a sophisticated signalling play, the question is whether it would have closed anyway. Usually the thing that closed it was a number, stated plainly, that somebody could check.
What this ensemble cannot see
All three models take the counterparty’s depth as the variable. None of them measures yours.
A founder who is themselves at level 1, best responding to an imagined investor who does not exist, will not be rescued by any of this. Worse, the depth model gives you a comfortable explanation for every rejection, and a model that explains all your failures is not a model. That is the trap this lens sets, and it is why the transparency card above names it.
There is also the auction problem, which is not a small exception. Fundraising has the shape of a common-value contest, and that is the one environment where the depth model has been tested against equilibrium and lost.2 If your situation is genuinely competitive bidding, weight this lens down and reach for the bidding models instead.
The one action that survives the ignorance: before the next pitch, hand your deck to one person outside your company and ask them a single question. What is this company claiming, in one sentence? If their answer requires a second sentence, you were writing for depth three, and you have until the meeting to fix it.
Who has to move
This changes nothing unless whoever writes the deck reads it, and in most companies the founder writes the deck and is also the person most certain that their meaning is obvious. The cheapest first test costs one conversation with an outsider and takes ten minutes. If the sentence comes back wrong, everything above is worth doing. If it comes back right, you have a different problem and should stop reading about signalling.
Sources and notes
- Jeffrey Carpenter and Andrea Robbett, Game Theory and Behavior, MIT Press. Level-k reasoning and the beauty contest are chapter 27, including Nagel’s finding that guesses cluster on the level-1 and level-2 predictions and the observation that roughly ninety per cent of players in the surveyed matrix games split between those two levels. The cognitive hierarchy model of Camerer, Ho and Chong is in the same chapter. Verifiable disclosure, unravelling and the Los Angeles restaurant hygiene card case are in chapter 12, along with the point that unravelling does not begin without a credible means of sharing information.
- Stafford Beer, Cybernetics and Management, English Universities Press, 1959. The black box strategy, the child and the cot, the entropy-of-selection argument for how to search, and the criticism of cause-effect cost analysis as a homomorphism in which causal relationships are assumed but do not hold, are all in the chapter on the Black Box. Note: the copy consulted is an image-only scan and was read via optical character recognition, so it is cited qualitatively and no figure is quoted from it.
- Itzhak Rasooly, Going, Going, Wrong: a test of the level-k (and cognitive hierarchy) models of bidding, University of Oxford Department of Economics. https://www.economics.ox.ac.uk/sites/default/files/economics/documents/media/going_going_wrong_a_test_of_the_level_k_itzhakrasooly.pdf. The abstract states that when plausibly calibrated the level-k model substantially under-predicts observed bids and is clearly outperformed by equilibrium, that fitting it produces implausibly high estimated levels bearing no relation to levels inferred from a game known to trigger level-k reasoning, and that subjects almost never appeal to iterated reasoning when asked to explain how they bid.
- Ginger Zhe Jin and Phillip Leslie, The Effect of Information on Product Quality: Evidence from Restaurant Hygiene Grade Cards, Quarterly Journal of Economics 118(2), 2003. Open-access copy: https://drum.lib.umd.edu/bitstreams/1fce72fc-166b-49ba-8167-fc69cd0b13a6/download. The abstract states that the grade cards cause restaurant health inspection scores to increase, consumer demand to become sensitive to changes in restaurants’ hygiene quality, and the number of foodborne illness hospitalisations to decrease.
- Rosemarie Nagel, Unraveling in Guessing Games: An Experimental Study, American Economic Review 85(5), 1995, pages 1313 to 1326. The American Economic Review copy is paywalled; an accessible mirror carrying the game design, the level-k classification and the pooled results is Nagel’s own slide set hosted at Stanford: https://web.stanford.edu/~niederle/GuessingGames.pdf. Cited for the design and the clustering result, not for a specific figure.
- Colin F. Camerer, Teck-Hua Ho and Juin-Kuan Chong, A Cognitive Hierarchy Model of Games, Quarterly Journal of Economics 119(3), 2004, pages 861 to 898. Accessible copy: https://www.csc2.ncsu.edu/faculty/mpsingh/local/Social/f25/wrap/readings/Camerer+Ho+Chong-cognitive-hierarchy-of-games-2004.pdf. The one-parameter formulation, in which the frequency of players using k steps of reasoning follows a Poisson distribution, is the model referred to here.
A note on why the failure is in the body and not the footnotes. A framework that only reports the settings where its models win is marketing. The level-k lens fails in auctions, fundraising is auction-shaped, and a founder reading this deserves to know that before they use it in a round rather than after.
Joshua Agonya Pi’Rwot, Founder.