Diligence re-derives your numbers from raw records, and the analyst doing it carries an asymmetric penalty for missing the lower version. Most headline metrics hold enough definitional slack that the re-derivation lands somewhere else. Retention is a ratio computed inside a closed cohort, so the slack is narrow and the number that comes back resembles the number you sent.
That is the mechanic. A level is a stock, and a stock is a decision about which rows to count. A retention rate is a transition inside a population fixed before anyone had reason to move it.
Build the company around the number that holds. Then build the data room so a stranger can rebuild it without you in the call.
Why these three models
This piece runs the Wire Model: score the features of the decision, route to a small ensemble of formal models, force the ensemble to produce dated actions. The scores that mattered:
- Cohort state-transition structure (0.9). The underlying object is a population moving between states. Every headline number is a sum over those transitions, and the sum can be re-taken where the transition cannot.
- Measurement degrees of freedom (0.85). One event log supports many defensible definitions of the same metric. Diligence collapses them to one, and you do not pick which.
- Strategic actors with asymmetric information (0.7). You choose the definition. They choose the re-cut. Only one of you has seen the raw file.
- Cognitive and affective distortion (0.6). Both sides over-read small cohorts and under-read denominators.
- Regime-break risk (0.5). Usage pricing and machine-initiated purchasing are changing what a retained account refers to.
- Exposure topology (0.15). Scored and rejected. Network position does not change what the ledger says, so no network model enters.
That routes to Markov cohort-transition (why retention holds), mechanism design under strategic reporting (why everything else decomposes), and behavioral (how both sides misread the table). They span three outcome types, cycle-regime, equilibrium and noise, so the ensemble’s mistakes do not stack. Growth accounting scored well as a fourth and was cut: its one lever, where the marginal dollar goes, falls straight out of the Markov steady state.
The framework: every metric has a re-cut spread
Carry this handle into the meeting. A metric’s re-cut spread is how far it moves when a stranger rebuilds it from your raw records under a different but defensible definition. Wide spread, it decomposes. Narrow spread, it survives.
1. The transition lens: retention is a rate inside a frozen population
Model the customer base as states and transitions. Active, dormant, churned, resurrected. Every number on your deck is a sum over the transitions inside a window.
The sums are re-takeable. Which orders count when a third are cancelled on delivery. Which currency, and at which rate on which date. Whether a reactivated buyer counts as new. Whether a non-binding commitment counts as revenue. Each choice is defensible, each moves the total, and the analyst makes all of them.
A transition rate refuses most of that. Fix the cohort at t0, then count how many of exactly those customers transacted again by t1. The denominator was set before the incentive to move it existed, so mix shift cannot touch it.
The steady state falls out of the same arithmetic. With constant acquisition and constant churn, the base converges to acquisition divided by churn. Doubling acquisition doubles the base once. Halving churn doubles the base and keeps doubling every cohort after it. Across the five firms Gupta and colleagues valued, retention elasticity ran between 3 and 7, while acquisition-cost elasticity ran between 0.02 and 0.3.2 Those ranges are nowhere near each other. Spend accordingly.
One trap here will get you caught. Cohort retention rates typically rise with tenure, and the cause is sorting rather than loyalty: high-churn customers leave first, so survivors re-weight the average upward while every individual’s churn propensity stays flat.1 Never present a rising curve as evidence the product got stickier. An analyst reads that as a mixture, re-cuts it into segments, and the segments say something you did not.
Africa standardised on the transition years ago, because the stock was uninformative. Mobile money reached 2.3 billion registered accounts in 2025 against 593 million 30-day active accounts, an activity rate of 25.7%.3 A level is an opinion about which rows to count. A rate inside a closed cohort is not. The industry running on agent floats and USSD menus reports the rate, and the rate is the only account number anyone quotes twice.
Markov cohort-transition, the structure lens
Assumes: a customer can be identified across two purchases, and the cohort is a closed set.
Fits because: cohort state-transition structure scored 0.9.
Breaks when: identity is broken. Cash on delivery, shared handsets, a new SIM after every promo, one agent wallet serving forty people. Then the cohort is an artifact and retention inherits the identity error.
Counteracts: the belief that a bigger top-line number is a stronger claim.
May reinforce: optimising for repeat transactions of any size, including the unprofitable ones.
2. The reporting lens: metrics decompose because they are paid to
Campbell stated the law in 1976 and it has not needed amendment: “the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures.”4 His illustration was Chicago voting statistics against census data. Votes decided jobs, money and power, so votes were attacked. Census numbers decided nothing, so census numbers stayed clean.
Your growth number decides money. Treat it as a vote count.
The mechanism has a formal statement. Give an agent several tasks and measure only some, and incentive pay reallocates effort towards the measured ones.5 Point a company at growth and the organisation produces growth-shaped artifacts. Discounted first orders. Reactivation campaigns booked as acquisition. Pilot contracts counted at their ceiling. Non-binding intent recorded as committed revenue.
Here is where that ends. In one case the SEC brought, a private company’s reported annual recurring revenue was built by inflating deal values and booking uncommitted amounts from non-binding agreements as guaranteed future payments, with invoices fabricated to match. The board’s internal investigation revised the valuation from roughly $1.1 billion to roughly $300 million, and around 70% of principal went back to Series B and C investors.6
The African version runs through the distribution channel, because that is where the incentives sit. Jumia disclosed improper orders placed and subsequently cancelled, including through its JForce agent network, worth approximately 2% of GMV in 2018 and approximately 4% in the first quarter of 2019, plus a separate commission-manipulation scheme worth about 1%.7 The agent network that delivered distribution delivered the corruption path. Agents behaved as a commissioned channel measured on order volume always behaves. Design the measurement accordingly.
Definitions get engineered too. Groupon took a metric into its IPO filing that stripped out online marketing, which is to say the cost of acquiring the subscribers the metric was counting. It did not survive the amended filing.8
So the regulator built a mechanism, and it is the cheapest thing here to copy. Where a metric is material, the guidance expects a clear definition and calculation, why it is useful to investors, and how management uses it. Change the method and you disclose the difference, the reason, the effect on figures already reported, and you consider recasting prior periods.9 Run that on yourself at pre-seed. One definitions file with a change log does most of the work.
Which gives the criterion. A metric has a narrow re-cut spread when a stranger can rebuild it from records you do not control, and when the cheapest way to move it is to do the real thing. Retention scores well on both. Settlement files and reorder records sit with your processor and your distributor, and the cheapest route to a higher month-three number is a customer who came back. GMV scores badly on both.
Mechanism design, the reporting lens
Assumes: reporting is strategic, and definitional discretion is worth money to whoever holds it.
Fits because: measurement degrees of freedom scored 0.85, asymmetric information 0.7.
Breaks when: retention becomes the promoted metric. Then Campbell applies to retention: redefine active as any app open, bury resurrections in the retained bucket, make cancellation a phone call.
Counteracts: the assumption that a number is true because your own system produced it.
May reinforce: compliance theatre, and a well-documented data room measuring the wrong thing.
3. The reading lens: both sides misread the same table
People treat a small sample as representative of the population it came from, and trained researchers do it as reliably as everyone else.11 A month-six retention figure computed on 40 customers carries an interval wide enough to hold both your thesis and its opposite. Founders present it as a fact. Investors receive it as one.
Two more failures do the rest of the damage. Survivorship: the table shows customers your system managed to record, and the ones who never activated never entered the denominator. Denominator drift: you compute retention on customers who placed a second order, the analyst computes it on everyone who placed a first, and that gap is the entire conversation.
The fix is procedural. Print n in every cell. Put the denominator definition above the table rather than in a footnote. Lead with the cohort you would rather not show. Where a customer can churn and return under a new number, say so before someone says it for you.
Behavioral, the reading lens
Assumes: misreading is systematic and directional, so it can be pre-empted by format.
Fits because: cognitive and affective distortion scored 0.6.
Breaks when: the reader is a modeller with the raw file. Then format stops mattering and only the data does.
Counteracts: confident storytelling over 40 customers.
May reinforce: paralysis, and waiting for statistical significance you will not have until Series A.
GEER: the levers, cheapest first
Four channels set the re-cut spread: identity (can a customer be tracked across two events), definition (what counts), reconstructibility (can a stranger rebuild it), behaviour (does the number move for real reasons). Pull the cheap and reversible first.
- Write the definitions page. Active, churned, resurrected, cohort, period, currency. Hits definition. Costs an afternoon.
- Print n on every cell. Hits reading, costs nothing, removes the first objection an analyst raises.
- Produce the raw export. Customer id, timestamp, amount, currency, channel, status. Then rebuild your reported top-line from it alone. Hits reconstructibility. Costs a day.
- Re-cut yourself, adversarially. Recompute the headline number under the three definitions an analyst would reasonably pick. Report the worst first. Hits everything. Costs a week.
- Get one series attested externally. Processor settlement file, mobile money operator statement, distributor reorder record. Hits reconstructibility. Costs weeks.
- Fix the identity layer. Make a returning customer resolve to the same record. Hits identity. Costs months, and no lever above it works without it.
No-lever flag: if you cannot link a customer across two purchases, you have no retention metric to defend and no presentation creates one. That is instrumentation, it costs a quarter, and it belongs in the plan rather than the deck.
RADAR: the portfolio, dated
DO NOW, by T+3 days. Reversible, dominant across every scenario.
- Publish the definitions page and attach it to every metrics deck from now on.
- Produce the raw event export and rebuild your reported headline number from it. Log the difference.
- Recompute retention on the widest defensible denominator, everyone who transacted once. Lead with that.
- Add n to every cohort cell.
HEDGE, by T+14. Cheap insurance against a re-cut you did not run.
- Get one retention series attested by a party you do not control.
- Version-control the definitions with a change log, including effects on figures already reported.
- Name the cohort that looks worst, write the two-sentence explanation, and file it before anyone asks.
DEFER AND TRIGGER. Irreversible, so pre-commit the observable trigger now.
- Defer: rebuilding the warehouse, restating publicly, changing your headline metric while a raise is open.
- Trigger to proceed: two independent re-cuts land within 10% of your reported number by T+14. The metric is load-bearing. Take it into the room.
- Counter-trigger: any re-cut moves the number more than 25%, or the export takes longer than 48 hours. Stop the process, fix instrumentation, and tell the investors already in diligence before they find it. They will.
If you are the one doing the re-cutting. DO NOW: request the raw export and the definitions page before the first meeting, rather than in week three. HEDGE: re-cut every company two ways and record the spread as a data point about the team. DEFER: changing your screening model until you have a full cycle of measured spreads. For calibration, median net revenue retention across more than a thousand private B2B software companies sits near 103%, gross revenue retention near 91%.12 Treat anything far above that as a definitional question first.
CHAIN: what usually happens next
Match the reference class on structure. Any setting where a self-reported number allocates money and a third party later re-derives it from primary records runs the same machinery: non-GAAP metrics under regulatory review, school test scores under the law Campbell wrote. The base rate holds. The re-derived number comes back lower, by roughly the amount of discretion the reporter held.4, 9
Second-order consequence: analysts stop asking for metrics and start asking for exports. Daniel McCarthy and Peter Fader rebuilt Blue Apron’s customer economics from the cohort data in its own IPO prospectus and put six-month churn near 70%, in public, before the stock priced.10 No private information was used. If you cannot produce the raw export in 48 hours, you do not have a retention number, you have a claim about one.
Third-order: teams that instrument early clear diligence faster, which reads as quality and partly is. Subtract that counterfactual. Retention is carrying credit that belongs to the whole operation.
Matrix-break flag. Usage-based pricing and machine-initiated purchasing are pulling apart the identity between a retained account and retained revenue. A customer can be fully retained while their spend halves, and a software agent can transact for months without a login, so seat and logo retention stop measuring what they used to. Short run, revenue retention by cohort, in the currency you bank, still separates. Medium run, record the acting entity on every transaction, human or machine. The cohort definition you need in two years is absent from your schema today.
What this ensemble cannot see
Four gaps, and none of them are small.
Whether the retained customers are worth keeping. Retention is silent on margin. A cohort returning monthly at negative contribution retains beautifully and destroys the company on schedule.
The customer who never entered the log. Everyone who bounced before your instrumentation fired sits outside every cohort here, and they are usually your largest population.
Off-platform behaviour. WhatsApp reorders, cash top-ups at an agent, a distributor buying through a side channel to protect its margin. Real retention, invisible to the export, common in the markets where the export matters most.
Retention’s own corruption. The metric that decides money is the metric that gets managed, and retention is next in line. Nothing in the arithmetic protects it once it becomes the thing investors pay for.
Run the ensemble on the part you control. The raw export survives every version of that ignorance, because it is the one artifact that lets someone else check you and lets you check yourself. Produce it by T+3. Re-cut it against yourself by T+14. If your worst re-cut still clears the bar, take that number into the room and lead with it.
Sources and notes
- Fader, P. S., Hardie, B. G. S., Liu, Y., Davin, J., and Steenburgh, T. “‘How to Project Customer Retention’ Revisited: The Role of Duration Dependence.” Journal of Interactive Marketing 43, 2018, 1-16. Working version, January 2018: full text. Mirror at LBS Research Online. The sorting result is stated in the abstract: increasing cohort-level retention rates under the beta-geometric model are “purely due to cross-sectional heterogeneity; an individual customer’s propensity to churn does not change over time.”
- Gupta, S., Lehmann, D. R., and Stuart, J. A. “Valuing Customers.” Journal of Marketing Research 41(1), 2004, 7-18. Full text. Retention elasticity of 3 to 7, margin elasticity about 1, acquisition elasticity 0.02 to 0.3, estimated across five firms.
- GSMA, State of the Industry Report on Mobile Money, reported in the GSMA press release “Mobile Money accounted for $2 trillion in transactions in 2025.” Release. 2.3 billion registered accounts, 593 million 30-day active accounts, 25.7% activity rate.
- Campbell, D. T. “Assessing the Impact of Planned Social Change.” Occasional Paper Series #8, The Evaluation Center, Western Michigan University, December 1976; published as Evaluation and Program Planning 2(1), 1979, 67-90. Full text. The quoted law and the Chicago voting-versus-census illustration are in the section “Corrupting Effect of Quantitative Indicators.”
- Holmstrom, B., and Milgrom, P. “Multitask Principal-Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design.” Journal of Law, Economics, and Organization 7 (Special Issue), 1991, 24-52. Full text.
- SEC v. Manish Lachwani, complaint filed in the U.S. District Court for the Northern District of California, August 2021. Complaint; SEC press release 2021-164. Allegations regarding ARR inflation, fabricated and altered invoices, the revision from approximately $1.1 billion to approximately $300 million, and the return of approximately 70% of principal to Series B and C investors are as pleaded by the Commission.
- Jumia Technologies AG, Q2 2019 results, Form 6-K exhibit filed 21 August 2019, section “Sales Practices Review.” SEC EDGAR. Contemporaneous coverage: TechCabal.
- Groupon, Inc., Form S-1 filed 2 June 2011, which introduced Adjusted Consolidated Segment Operating Income, a measure excluding online marketing expense. Original filing. Its removal from the amended registration statement is reported by Forbes, 10 August 2011.
- U.S. Securities and Exchange Commission, “Commission Guidance on Management’s Discussion and Analysis of Financial Condition and Results of Operations,” Release No. 33-10751, 30 January 2020. Release. The three expected disclosures and the change-in-methodology requirements are in Section I.
- McCarthy, D., and Fader, P. S. “Customer-Based Corporate Valuation for Publicly Traded Non-Contractual Firms.” Journal of Marketing Research 55(5), 2018, 617-635. The Blue Apron reconstruction and the six-month churn estimate, drawn from metrics disclosed in the company’s own IPO prospectus, are described in Knowledge at Wharton. Note: that URL slug contains a word banned in FounderWise prose. It is reproduced verbatim as part of the record.
- Tversky, A., and Kahneman, D. “Belief in the Law of Small Numbers.” Psychological Bulletin 76(2), 1971, 105-110. Full text.
- SaaS Capital, “2026 Benchmarking Metrics for Bootstrapped SaaS Companies,” survey of more than 1,000 private B2B software companies. Report. Median net revenue retention 103%, median gross revenue retention 91%, for companies between $3M and $20M ARR.