FounderWiseDecisions, not feeds
← All articles FounderWise · Long-form

You are grading them on a signal you cannot read

You can see the result. You cannot see the effort. Every performance judgement you make is an inference through noise, and most management systems pretend otherwise.

28 Sep 2026 13 min read By Joshua Pi’Rwot
Share X LinkedIn

Two salespeople. One closed four deals this quarter and one closed none. You are about to promote the first and manage the second out.

You have not observed either of them working. You observed two outcomes and inferred two people, and the inference is doing far more work than you think.

Why these three models

The decision is what to conclude about a person from their results, and what to attach rewards to. The features that fire are an action you cannot watch directly, an outcome that depends on the action and on much else besides, and a sample too small to separate the two.

Three lenses. Monitoring noise produces a complex answer about why the outcome is an unreliable reading of the action. Mechanism design produces an equilibrium answer about what happens once people know what is being measured. Luck and skill produce a random answer about how many periods you need before a difference in results is evidence of a difference in people. The third is the one that decides whether the first two matter, and it is the one nobody runs.

1. The outcome is a reading, not the thing

Start with the structure rather than with the people. The person acts. The action combines with everything outside their control. You observe the combination. You never observe the action.

That is the standard hidden-action setting, and its central result is not about laziness. It is that when monitoring is imperfect, punishing bad outcomes necessarily punishes bad luck as well, because you cannot tell them apart from where you are standing. Any scheme strong enough to deter shirking also imposes cost on people who did everything right.1

The size of that problem is set by one thing: how much of the variation in the outcome comes from the action versus from everything else. In a support queue where the tickets are similar, most of the variation is the person, and outcome-based judgement is close to fair. In enterprise sales with six-month cycles and twelve deals a year, most of the variation is territory, timing, and which two accounts happened to have budget, and outcome-based judgement is close to a lottery you are calling a meritocracy.

The first question is therefore not how do I measure performance. It is how noisy is my measurement, and that is answerable. Compare the same role across people in the same period, and the same person across periods. If the spread within one person over time is as wide as the spread between people, you are reading noise.

Run that comparison with real numbers and it stops being abstract. Take one sales role over eight quarters. Rep A closed 5, 1, 3, 6, 2, 4, 1, 5. Rep B closed 3, 2, 4, 2, 3, 1, 4, 2. A’s own range runs from 1 to 6, and the gap between the two averages is under one deal. Pick any single quarter and you can make either of them look like the strong one. That is the whole calculation, and it runs in an afternoon on a spreadsheet you already keep.

One distinction the spread will not make for you. Noise is not the same as an unequal deal. A rep working Kampala industrial accounts and a rep working upcountry retail are not two draws from the same process, and no length of run makes them comparable. Ask first whether the territories are the same game. If they are not, compare each person against their own history and against what that territory did last year, never against each other.

2. People optimise against what you can see, not what you want

Once you accept that the observation is noisy, the design question follows: what should you attach consequences to.

The mechanism design answer is that you can only build on what is observable and verifiable. Intentions, effort and judgement are none of those things inside a company. Whatever you actually attach rewards to becomes the objective, regardless of what the objective was supposed to be, and this is not cynicism about people. It is the predictable response of a reasonable person to a stated rule.2

Which produces the common failure. A team measured on tickets closed closes easy tickets. A team measured on deals closed discounts. A team measured on lines of code writes more of them. In each case the metric was a proxy for something worth having, and the proxy was the only part anyone could see, so the proxy is what got produced.

The useful move is narrower than most advice suggests. Rather than searching for a metric that cannot be gamed, which does not exist, ask what a person could do to move this number that I would not want. Then check whether that action is cheaper than the real work. If it is, the metric is not usable as an incentive no matter how well it correlates with performance when nobody is being paid on it. Correlation measured before the incentive existed tells you nothing about behaviour after.

A metric measured for information and a metric attached to money are two different instruments, and only one of them changes what people do.

There is a practical consequence founders resist. Some things genuinely cannot be incentivised well and should be managed by observation, conversation and judgement instead. That is more expensive and it does not scale cleanly. It is also correct in exactly the cases where the noise is high, which is most of the early-stage company.

Managing by observation also has a scale at which it fails, and it is worth knowing roughly where yours is. It works while one person can hold the work of everyone they judge in their head, which in practice is somewhere under fifteen people and fewer if the work is varied. Past that, the judgement is not observation, it is recollection of the loudest events, which is a worse instrument than the noisy metric it replaced. The signal that you have crossed the line is that you start reviewing people from memory rather than from anything you watched that month.

And expect a request the design does not anticipate. Strong performers often ask to be paid on the number. They are not confused, they are pricing their own confidence, and a flat refusal reads as a lack of faith in them. The workable answer is to pay on the number with a floor, so the person carries the upside of the outcome without carrying the whole downside of a thin territory or a quiet quarter. That splits the risk rather than arguing about who owns it.

3. How many quarters before the difference is real

The third model is the one that turns the first two from philosophy into arithmetic.

Where outcomes depend substantially on chance, the number of observations needed before a difference in results reflects a difference in ability is much larger than intuition allows. Short sequences produce apparent gaps routinely with no underlying difference at all, and the more chance is involved, the longer the run required.3

Apply it to the two salespeople. One quarter, four deals versus none, in a market where the median rep closes two. That is not a sample. Under any reasonable noise assumption, that gap will occur regularly between two identically capable people. Promoting on it is not a mistake of kindness; it is a measurement error with a person attached.

There is a second error stacked on top of the first, and it runs in one direction. Faced with a result, people reach for an explanation in the person rather than in the circumstances, and they do it confidently on very little evidence.4 So the four-deal quarter does not merely get over-read. It gets converted into a claim about character, which is far harder to revise later than a claim about a number. The number was noisy, and you turned it into a description of who someone is.

The correction is not to wait forever. It is to change what you observe. Outcomes accumulate slowly and noisily. Actions accumulate quickly and cleanly. How many first meetings did they book, how many were with the right buyer, how much of the pipeline was self-generated. These are observable, largely within the person’s control, and available monthly rather than annually, which is what makes them usable evidence where the closed number is not.

Inputs are gameable too, and that failure is quieter because inputs are cheap to produce. A rep judged on first meetings will book meetings with whoever accepts. So attach a quality condition to each input rather than adding a second metric beside it: meetings with a named budget holder, not meetings; self-generated pipeline, not pipeline. And keep the list to three. A dozen inputs is a procedure, and a procedure is exactly the thing section two says people will follow instead of doing the work.

The horizon rule also cuts both ways, which is the half managers skip. If you will not dismiss somebody on one bad quarter, you cannot promote on one good one. Whatever run length you name, name it once, write it down, and apply it in both directions. A company that is patient with failure and impatient with success has not set a horizon, it has set a bias.

The rule that falls out is simple enough to apply. Judge people on inputs at short horizons and on outcomes only at long ones. Most companies do exactly the reverse, and then wonder why their performance system feels arbitrary to everyone inside it.

What the three say together

  • Estimate the noise. Is the spread within one person over time as wide as the spread between people. If yes, outcome-based judgement is not working.
  • List what a person could do to move each incentivised number that you would not want. If the cheapest route to the number is not the real work, do not attach money to it.
  • Set the horizon honestly. Name the number of periods you would need before a difference in outcomes is evidence, and stop making outcome-based calls before it.
  • Move short-horizon judgement onto observable inputs, and keep outcomes for the long horizon where they finally mean something.

Where they disagree

Mechanism design and the noise argument pull in opposite directions on how hard to tie reward to result.

Mechanism design says a weak link between reward and outcome invites shirking, because the person bears little of the consequence of their own choices. The noise argument says a strong link punishes bad luck and drives out good people who had a bad quarter through no fault of their own. There is no setting that satisfies both, and the standard treatment is explicit that the trade-off is real rather than an artefact of poor design.

The resolution is that the correct strength of the link is set by the noise, not by your management philosophy. High noise means weaker outcome-linkage and more input measurement, and you accept the shirking risk because the alternative punishes the wrong people. Low noise means the reverse. A founder who applies the same compensation philosophy to a support team and an enterprise sales team has answered a question that has two different answers.

What none of them contain

None of the three accounts for what the measurement does to the person being measured. All of them treat the employee as responding to incentives while remaining otherwise unchanged. In practice, being judged on a noisy signal produces its own effects, and being judged on inputs can feel like surveillance even when it is more accurate. The models are silent on how it feels, and how it feels determines who stays.

None of them handles the person whose contribution is largely to other people’s numbers. The engineer who unblocks three colleagues a week has an outcome, and it is recorded against someone else. Every scheme in this article will underrate them, and the only known fix is a judgement call made by someone close enough to see it, which is the thing all this machinery was built to avoid needing.

And none of them tells you what to do about the person you have already misjudged. The models are prospective. The repair is not a modelling problem.

The one action that survives the ignorance: take the role where you are most confident in your performance ranking and check the spread of the same person’s results across the last four periods. If one person’s own quarter-to-quarter range covers most of the gap you believe exists between your best and your worst, you have learned that your ranking is partly noise, and you have learned it in an afternoon without anyone being told they are being assessed.

Who has to move

This sits with whoever sets compensation, which in an early company is the founder and nobody else. The instinct when the performance system feels unfair is to add more metrics, which increases the number of gameable proxies without reducing the noise in any of them. The cheapest first test is one role, four periods, one spread calculation. It costs an afternoon and it is the only piece of evidence that tells you whether your current system is measuring people or weather.

Sources and notes

  1. Jeffrey Carpenter and Andrea Robbett, Game Theory and Behavior, MIT Press. The hidden-action problem, imperfect monitoring, and the trade-off between deterring shirking and imposing risk on the agent are developed in the chapters on asymmetric information and contracts. Used in section 1 for the central claim that any scheme strong enough to deter shirking under noisy monitoring necessarily also punishes bad luck.
  2. Jeffrey Carpenter and Andrea Robbett, Game Theory and Behavior, MIT Press. Mechanism design, incentive compatibility, and the requirement that a scheme be built on observable and verifiable quantities are treated alongside auctions and contract design. Used in section 2. The related failure, in which a measure ceases to be a good measure once it becomes a target, is treated in Donella H. Meadows, Thinking in Systems: A Primer, Chelsea Green Publishing, under seeking the wrong goal.
  3. Michael J. Mauboussin, The Success Equation: Untangling Skill and Luck in Business, Sports, and Investing, Harvard Business Review Press. The luck-skill continuum, and the consequence that the sample required to distinguish ability grows with the chance component, underpin section 3. The claim used here is directional. No specific number of quarters is asserted for any real sales role, because that number depends on the variance of your own pipeline and can only be computed from it.
  4. Sanjit S. Dhami, The Foundations of Behavioral Economic Analysis, Oxford University Press. Attribution of outcomes to disposition rather than to situation, and the small-sample confidence that accompanies it, sit in the part on bounded rationality. Cited for direction only, with no magnitude claimed, and used to support the framing in section 3 that the error runs toward reading people out of results.

A note on a number this article does not give. There is no universal ratio at which a metric becomes safe to pay on. It depends on how much of your outcome variance the person controls, which differs by role and by company and changes as you grow. What transfers is the spread calculation, and you can run it this week on data you already have.

Joshua Agonya Pi’Rwot, Founder.

Lock in your calls.

You’ve marked 0 of 5. Now choose how often you want the signals.

Step 1 · Pick your cadence

The DispatchWeekly · your Monday 5 callsFreealways

Step 2 · Where to send it

Personalize your BriefThe Brief

Tune every edition to the markets and industries you actually act on.

🔒 Unlock personalization — The Brief, $19.99/mo →
Free Dispatch forever · upgrade anytime · we never share your details.
Know where you stand?
The Dispatch tells you what changed. Knowing what to do about it is a different question, and it is the one FounderWise answers. Start with the free Traction Audit: 12 questions, about 3 minutes, scored out of 100.
Find out where you stand →

For teams, syndicates & programs

Recommended
Team
$15/seat · mo
Daily Brief for the whole team (min 3 seats).
  • Everyone on the same signal
  • Admin + shared watch-list
  • One invoice · ~25% off solo
Get Team →
Cohort Licence
$2,900/yr
Co-branded seats for one cohort, for accelerators, funds & programmes.
  • Up to 25 founder seats
  • Your logo, your cohort
  • The record of what your cohort committed to
Talk to us →
Pass the Dispatch on
Know a founder making these calls blind? Send them this week’s five — free, every Monday.

Decisions, not feeds. · Curated by Joshua Pi’Rwot · FounderWise · Free Audit · Store · parent of Business Growth Accelerator

Call committed. We’ll hold you to it.
Know where you stand →