A number you cannot explain is usually a customer doing a job you never built for. The reflex that destroys it is reconciliation: you reach for the story that puts the line back where you expected it, and the search stops there. Find the person behind the number before you explain the number.
Three questions decide what an anomaly is worth. Is it real. What position does it occupy. Will it spread.
Most teams get the first one wrong and never reach the other two.
Why these three models
This piece runs the Wire Model: score the features of the decision, route to a small ensemble of formal models, then make the ensemble produce dated actions. What the scoring turned up:
- Small-sample variance (0.85). Weekly product data is thin. An extreme reading is the single most likely thing to be produced by chance.
- Heterogeneous demand hidden by an aggregate (0.8). Dashboards report central tendency. Demand is spread across positions, and the spread is where the product lives.
- Threshold dynamics (0.7). A workaround either stays with the nine accounts doing it or tips into a segment, and the tipping depends on the shape of the distribution rather than its mean.
- Cognitive distortion (0.75). The reflex itself. Familiarity with what your product is for blocks you from seeing what it is being used for.
- Deep uncertainty (0.55). You cannot see willingness to pay, and you cannot see the users who wanted this and left.
Routing gives three models. The luck-skill continuum asks whether the reading survives (random). Spatial choice asks what position it occupies (equilibrium). The threshold model asks whether it spreads (complex). Three outcome types, so the errors point different ways.
Cognitive distortion scored high enough for a fourth card and does not get one. Its only useful lever, writing down what you expect before you look, already sits inside the first model’s rerun test, and the evidence for the reflex sits inside the second model’s section. A card that repeats two levers you have already been given is overhead. Three it is.
The framework: an anomaly is a customer standing somewhere you did not build
1. The filter: most anomalies are variance, and the survivors are not
Put the metric on a continuum with pure skill at one end and pure luck at the other. Where an outcome sits on that line tells you how fast an extreme reading reverts and how large a sample you need before the reading means anything. Outcomes loaded with luck revert to the mean faster, and a small number of observations makes it very hard to separate the two.1 Your weekly funnel is closer to the luck end than your board deck implies.
Microsoft’s experimentation team wrote down what this feels like from the inside. An experiment produced very surprising results, metrics unrelated to the change moved in unexpected directions, the effects were highly statistically significant, and when the team reran the experiment many of the effects disappeared. This happened often enough that they stopped treating it as a one-off and went looking for root causes.2 The same paper reports that most suspected primacy and novelty effects, the ones every product team narrates as users getting used to the new thing, are not real. They are a statistical artefact.2
So the first move on any anomaly is procedural, not analytical. Rerun the window. Pull the same query on a fresh cohort and a clean date range. Check the tracking change log before you check the chart. Most of what looks interesting dies here, and that is the point of doing it first.
The anomaly you can explain in one sentence is the one you have stopped investigating.
Luck-skill continuum, the filter lens
Assumes: the metric has a stable underlying mean, so extremes revert at a rate you can estimate.
Fits because: small-sample variance scored 0.85.
Breaks when: the process itself changed. Then reversion never comes, and you have discarded a real shift as noise.
Counteracts: narrating a two-week wobble as a trend.
May reinforce: reflexive dismissal, which is the exact failure this article is about.
2. The position: your dashboard reports an average, your customers occupy a space
Hotelling modelled buyers as points spread along a line and sellers as points that choose where to stand. His conclusion was that competitors converge toward the middle and the market ends up under-served at the edges. He put it plainly at the end of the paper: some factories make cheap shoes and others make expensive shoes, but all the shoes are too much alike, and cider is too homogeneous.3 Read that as a product statement. Every team building to the average of its own dashboard is walking toward the same crowded point.
A dashboard reports the average. Nobody is the average.
An anomaly is a customer standing away from that point, loudly enough that your instrumentation caught them. Von Hippel named these people lead users: their present strong needs become general in the market months or years later. He also explained why you cannot see them. Users steeped in the present are unlikely to generate concepts that conflict with the familiar, and the more recently an object has been used in a familiar way, the harder people find it to use that object in a new way.4 You built the product. You are the most functionally fixed person in the room.
The commercial gap is measurable. In a natural experiment at 3M, ideas produced by the lead user process carried projected year-five annual sales of $146 million, more than eight times the forecast for traditional projects run at the same time in the same divisions.5
The cleanest worked example of this is East African, and it happened on a screen. During the 2005 M-PESA pilot, built to let microfinance clients repay loans by phone, the team watched transaction patterns on their web screens and could not account for what they were seeing. They sent researchers to find out. What came back was a list: people paying each other for trades between businesses, larger traders using the float as an overnight safe because banks closed before the agent shops did, people depositing cash in one pilot area and withdrawing it hours later in another, one woman sending money to a husband who had been robbed so he could pay his bus fare home.6
None of that was loan repayment. The team did not reconcile it. A workshop with Safaricom’s commercial team turned the anomaly into the proposition Send Money Home, and the launch product was cut back to three things: deposit and withdraw cash at an agent, send money person to person, buy airtime.6 The feature the product was built for did not make the launch.
Your version of this is smaller and sitting in your own data right now. Retailers using your inventory app as a WhatsApp order log. Distributors screenshotting your invoice page because they need something to send a bank. Agents topping up float at 5am, hours before you thought anyone was awake.
Spatial choice, the position lens
Assumes: needs are spread across a space, and a product is a point chosen inside it.
Fits because: heterogeneous demand hidden by an aggregate scored 0.8.
Breaks when: the outlier is one determined customer rather than a thin edge of a real distribution. Then you have built a bespoke tool and called it a segment.
Counteracts: optimising toward the mean of your own users.
May reinforce: chasing every edge case into a product with no centre.
3. The spread: the shape of the distribution decides, not its average
Granovetter’s threshold model gives each person a trigger point, the number of others who must act before they do. Start with 100 people whose thresholds run 0, 1, 2 and so on to 99 and everybody eventually acts. Remove the person with threshold 1, replace them with a second person at threshold 2, and the process halts after one actor. The two crowds are essentially identical by any normal description. His conclusion is the one to keep: it is hazardous to infer individual dispositions from aggregate outcomes.7
That is a warning written for anyone reading a dashboard. Two products with the same weekly average can be one step apart in whether a workaround propagates. The nine accounts doing the strange thing are only interesting if there is a next tier of accounts whose trigger point is nine.
So ask a structural question rather than a volume question. What has to be true for the tenth account to start. Usually it is one of three things: they have to see someone else do it, the workaround has to get cheaper, or someone has to tell them it is allowed. All three are things you can supply deliberately.
M-PESA is again the useful measuring stick, because the anomaly did tip. From a pilot of a few hundred microfinance clients, the service reached roughly 65% of Kenyan households by the end of 2009.8
Threshold model, the spread lens
Assumes: adoption depends on the distribution of individual trigger points, which is at least partly observable.
Fits because: threshold dynamics scored 0.7.
Breaks when: the behaviour is private. If nobody can see anyone else doing it, the cascade mechanism never engages and the distribution is irrelevant.
Counteracts: reading a flat average as a flat market.
May reinforce: waiting for a tip that was never in the distribution to begin with.
GEER: the levers, cheapest first
Four channels carry the value of an anomaly: detection (does it reach you at all), attribution (do you know which accounts), replication (does it survive a rerun), packaging (can a stranger buy it). Pull the cheap reversible levers before the expensive irreversible ones.
- Log the prediction before you open the dashboard. One line, in writing, on what you expect each number to say. An anomaly only exists against a recorded expectation. Ten minutes a week. Hits detection.
- Look at the distribution, not the mean. Split by cohort, by geography, by device, by day of week. Hits detection. Costs an afternoon.
- Name the accounts. Turn the spike into a list of identities. If you cannot, stop here, because you have an instrumentation problem rather than a product signal. Hits attribution.
- Rerun on a clean window. Fresh dates, fresh cohort, tracking change log checked. Hits replication. Costs a day.
- Call five of them. Ask what they were trying to get done, and what they used before you existed. Hits attribution and packaging. Costs a week.
- Instrument the workaround. Add an event for the thing they are doing that you never built. Hits detection permanently. Costs a sprint.
- Sell the manual version. Deliver it by hand to ten accounts and charge for it. Hits packaging. Costs a month.
- Reposition the product around it. Hits everything. Costs quarters and is hard to reverse.
No-lever flag: if the anomaly cannot be traced to identifiable accounts and cannot be reproduced on a clean window, no amount of discussion converts it into a decision. Fix the instrumentation and put the question back on the calendar.
RADAR: the portfolio, dated
DO NOW, by T+3 days. Reversible, and worth doing whichever way the anomaly turns out.
- Start the prediction log. Write what you expect before every weekly review.
- Export the account list behind the anomaly, with names, plan and signup date.
- Rerun the query on a clean window and check what shipped or changed in tracking during the original one.
- Book five calls with accounts on that list. Ask what they were doing before your product existed.
HEDGE, by T+14. Bounded spend that pays off only if the anomaly is real.
- Ship an event that logs the workaround explicitly, so the next occurrence is measured rather than noticed.
- Deliver the workaround manually to ten accounts, by hand, without building anything.
- Quote a price to three of them. Usage tells you about interest. Only a price tells you about a market.
DEFER AND TRIGGER. Irreversible, so wait, and pre-commit the observable trigger today.
- Defer: roadmap surgery, renaming the product, rebuilding onboarding around the new behaviour, telling investors you have found a second market.
- Trigger to commit: the behaviour reproduces on a clean window, at least ten accounts you never contacted do it unprompted, and two of the manual deliveries convert to paid inside 14 days. Then rebuild.
- Counter-trigger: by T+28 the cohort has not grown beyond the original accounts and nobody pays. File the finding with the account list attached and move on. You found one customer with an unusual workflow. Serve them by hand.
Reading someone else’s anomaly. If you are the investor or the board member being shown the chart, three questions do most of the work. Was the expectation written down before the data arrived. Does it survive a rerun. Which named accounts, and has anyone spoken to them. A founder who can answer all three is running a process. A founder who answers none is telling you a story that was assembled after the fact.
CHAIN: what usually happens next
Structure decides the comparison set here, and the structure is shared by every organisation that ran a process, got a reading it did not order, and had to choose whether to chase it. Drug trials logging an unintended effect. Industrial firms discovering users who had already modified the equipment. Experimentation teams at scale explaining puzzling outcomes. Your category is irrelevant to this list.
The base rate from that class splits cleanly. Most surprising readings do not survive contact with a rerun.2 The few that do tend to be disproportionately large. At 3M, every funded lead user idea was for a major product line, while among the funded ideas from conventional methods one was a major product line and 41 were incremental.5 High mortality, heavy tail. That combination argues for a cheap standing filter rather than a heroic annual investigation.
Second-order effect: teams that install the filter start pivoting more, not less. In a randomised trial of 116 Italian startups, founders trained to build explicit frameworks and test their hypotheses the way a scientist would performed better, were more likely to pivot to a different idea, and were no more likely to drop out than the control group.9 Rigour raises the pivot rate because it makes a negative result legible instead of embarrassing.
Third-order effect: the accounts you call become the people who tell you the next thing first. That is a compounding asset and it is why the calls matter more than the query.
One correction before you credit the anomaly with anything. Ask what those same accounts would have done in the same period with no change at all, and take that off the number. Seasonality, a partner’s campaign, a competitor’s outage and a public holiday all produce charts that look like discovery.
Matrix-break flag. Automated anomaly detection now writes the explanation for you, fluently, in the dashboard, seconds after the reading appears. The reconciliation reflex has been productised and it is faster than you are. The models above assume a human pauses at the strange number. Where that pause has been automated away, the pause has to be reinstated deliberately, as a rule about which anomalies a person must look at before any generated commentary is read.
What this ensemble cannot see
Start with the largest gap. None of these models can tell you which tail behaviour is a market and which is one stubborn customer with an unusual workflow, because that distinction only resolves after you have priced it. Usage data is silent on willingness to pay.
The second gap is worse because it leaves no trace. Everyone who wanted this and left before your instrumentation existed is absent from every chart you own. Your anomaly is drawn only from survivors, and the survivors are the people your product almost served.
The third is you. The team’s incentive to reconcile is real: an unexplained number means a reopened roadmap, a delayed launch, an uncomfortable board slide. That incentive does not appear in any model here, and it is the mechanism that closes most of these investigations before they start.
None of that changes what to do on Monday. Write down what you expect before you look, and call five of the people who did the thing you cannot explain. Do both this week. The rest of the portfolio waits on the rerun.
Sources and notes
- Mauboussin, M. J. “Untangling Skill and Luck.” Legg Mason Capital Management, 15 July 2010. The skill-luck continuum, the rate of reversion to the mean, and the sample-size argument are on pages 3 to 5 and 12. Full text.
- Kohavi, R., Deng, A., Frasca, B., Longbotham, R., Walker, T., and Xu, Y. “Trustworthy Online Controlled Experiments: Five Puzzling Outcomes Explained.” Proceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD ’12), 786-794. The rerun that made the effects disappear is item 5 of the paper’s summary of puzzling outcomes; the finding that most suspected primacy and novelty effects are a statistical artefact is in section 3.4. Full text.
- Hotelling, H. “Stability in Competition.” The Economic Journal 39(153), 1929, 41-57. The shoes and cider passage is the closing paragraph, page 57. Full text.
- Von Hippel, E. “Lead Users: A Source of Novel Product Concepts.” Management Science 32(7), 1986, 791-805. Definition of lead users and the problem-solving evidence on familiarity blocking novel use are on pages 1 to 3 of the author’s manuscript. Author PDF.
- Lilien, G. L., Morrison, P. D., Searls, K., Sonnack, M., and von Hippel, E. “Performance Assessment of the Lead User Idea-Generation Process for New Product Development.” Management Science 48(8), 2002, 1042-1059. The $146 million year-five projection and the eight-times comparison are in the abstract; the major-product-line counts are in the results section. Author PDF.
- Hughes, N., and Lonie, S. “M-PESA: Mobile Money for the ‘Unbanked’. Turning Cellphones into 24-Hour Tellers in Kenya.” Innovations: Technology, Governance, Globalization 2(1-2), 2007, 63-81. The observed pilot behaviours and the decision to send researchers are on pages 75 to 76; the Send Money Home proposition and the three-feature launch product are on page 77. MIT Press bot-blocks direct fetches, so this links the identical article as mirrored by GSMA.
- Granovetter, M. “Threshold Models of Collective Behavior.” American Journal of Sociology 83(6), 1978, 1420-1443. The 100-person example and the warning against inferring individual dispositions from aggregate outcomes are on pages 1422 to 1424. Full text.
- Jack, W., and Suri, T. “Mobile Money: The Economics of M-PESA.” NBER Working Paper 16721, 2011. The figure of roughly 65% of Kenyan households by the end of 2009 is on page 2. Working paper PDF.
- Camuffo, A., Cordova, A., Gambardella, A., and Spina, C. “A Scientific Approach to Entrepreneurial Decision Making: Evidence from a Randomized Control Trial.” Management Science 66(2), 2020, 564-586. Sample of 116 Italian startups; treated founders performed better, were more likely to pivot, and were not more likely to drop out. Accepted manuscript.