Team 30 per cent, market 25, traction 25, product 20. Somebody wrote those numbers in a meeting and nobody has questioned them since. They now decide which companies you back, which candidates you hire, and which suppliers you qualify.
Ask where they came from and the honest answer is that they were asserted. A weighted score with no consistency check is a number with no standing.
Why these three models
The decision is whether to keep using a weighted scorecard, and if so how to derive its weights. The features that fire are a multi-criteria judgement, several people supplying inputs, and a systematic reasoning failure that nobody in the room can detect unaided.
Three lenses. Pairwise weighting produces an equilibrium answer about how to elicit weights and test them. Aggregation produces a random answer about whether combining several people’s scores helps or merely launders one opinion. Feedback misperception produces a complex answer about why the incoherence goes unnoticed. The first two disagree in a useful place: one is about a single decider’s coherence, the other is about what happens when you average across people, and running the first on the output of the second is a specific and common error.
1. Deriving weights instead of asserting them
The alternative to inventing percentages is to compare criteria against each other, two at a time, and let the arithmetic produce the weights.
Saaty’s procedure is straightforward. For each pair of criteria, state how much more important one is than the other on a ratio scale of one to nine, filling a reciprocal matrix. The normalised principal eigenvector of that matrix gives the weights. There is a simple approximation that avoids the eigenvector entirely: normalise each column, average across the rows, and use the resulting column.1
The step that matters is the one everybody skips. Compute the consistency index as the largest eigenvalue minus the number of criteria, divided by that number minus one. Divide it by the average consistency of a random matrix of the same size. That gives the consistency ratio, and Saaty’s guidance is to revise if the ratio is considerably higher than ten per cent.1
Read what that test actually does, because it is not a test of the world. It checks whether your own stated preferences contradict each other. If you say team matters three times more than market, and market twice as much as traction, then you have committed to team mattering six times more than traction. Say something else and the arithmetic catches you.
Two constraints come with it. Keep the matrix at seven criteria or fewer; Saaty notes one rarely needs to go beyond seven by seven in order to keep consistency reasonably high.1 And use the ideal-mode normalisation, dividing by the largest component rather than the sum, because the original form suffers rank reversal: adding an irrelevant alternative can flip the ranking of the ones already there.
Pairwise weighting, the coherence lens
- Assumes: you can compare criteria two at a time on a ratio scale, and that your comparisons should be internally consistent.
- Fits because: you are already weighting criteria, just without deriving or testing the weights.
- Breaks when: used in original normalisation with a changing alternative set, where rank reversal makes the output arbitrary.
- Evidence: grade C plus. Widely used, and the rank-reversal critique is real and was never fully resolved. Ideal mode fixes it; the original does not.
- Counteracts: weights invented in a meeting and never revisited.
- May reinforce: false precision, because a consistent set of weights can still be consistently wrong.
2. What happens when several people score
The second lens is about the step most teams take next, and it is where the method quietly breaks.
Aggregating several people’s estimates works when their errors are partly independent. Diversity subtracts from collective error, and the combination beats the average member. That is a real and useful mechanism, and it is the reason a panel is worth having at all.
The condition is the whole thing. Where the scorers talk before scoring, or where one senior voice anchors the room, the independence is gone and so is the benefit. What you have then is one opinion recorded four times, and averaging it produces a number with four signatures and one source.
Now combine that with the first model and the error becomes specific. Running a consistency check on an averaged committee matrix tests the coherence of a construct that is nobody’s actual preferences. Worse, averaging tends to smooth exactly the contradictions the check exists to catch, so a committee matrix will often pass the test that each individual matrix would fail.
The consistency check is for one decider. Run it before you average, never after.
Crowd aggregation, the combination lens
- Assumes: individual estimates carry partly independent errors, so combining them cancels some.
- Fits because: more than one person contributes to the score.
- Breaks when: scorers observe each other first, which collapses the independence that made aggregation work.
- Evidence: grade B plus. The mechanism is well established and the independence condition is routinely violated in practice.
- Counteracts: treating a room’s average as if it were a crowd’s.
- May reinforce: avoiding discussion entirely, which throws away genuine information transfer.
3. Why nobody notices
The third lens explains why an incoherent scorecard survives years of intelligent people using it.
People are poor at inferring the consequences of their own stated relationships, and the failure is robust. In controlled experiments, highly educated adults cannot infer the behaviour of even simple systems, and the persistent poor performance is not attributable to an inability to interpret graphs, to contextual knowledge, to motivation, or to cognitive capacity.2
Applied here, that means the transitive implication of your weights is not something you can check by thinking harder about it. Three pairwise judgements imply a fourth, and no amount of care in the room surfaces the contradiction, because the arithmetic is doing work your intuition does not perform.
Which is the argument for the procedure rather than for better judgement. The consistency ratio is not an insult to the people scoring. It is a calculation they were never going to do in their heads, and its whole value is that it fires on people who are being careful.
Feedback misperception, the blind-spot lens
- Assumes: people cannot reliably infer the implications of relationships they themselves stated.
- Fits because: the incoherence has survived repeated use by competent people.
- Breaks when: the criteria set is tiny, where transitivity is checkable by inspection.
- Evidence: grade A. Replicated widely and robust to experience, incentives and market institutions.
- Counteracts: the belief that a careful team would have spotted it.
- May reinforce: outsourcing judgement to arithmetic that is itself resting on invented inputs.
The levers, cheapest first
- Count your criteria. More than seven and the comparisons become unreliable. Cut before you weight.
- Run the pairwise comparison alone, first. One person, one matrix, before any group conversation. It takes twenty minutes for five criteria.
- Compute the ratio and act on it. Above roughly ten per cent, revise the comparisons rather than shipping the weights.
- Score independently, then combine. Written scores submitted before discussion, not after. This is a calendar change, not a methodology change.
- Use ideal-mode normalisation. Divide by the largest component. It costs nothing and it removes rank reversal.
- Or drop the scorecard. If nobody will run the check, the weights are decoration, and a stated qualitative judgement is more honest than a number nobody validated.
What to do before the next committee
Do now, sized at half an hour, effect immediate. Take your existing scorecard and run the pairwise comparison on its criteria yourself, alone, then compute the ratio. Reversible, free, and dominant across every scenario about whether the weights are sound.
Hedge, where the premium is the whole loss, live before the next decision. Ask everyone to submit scores in writing before the meeting rather than during it. If the room was already independent you have added one email to the process, and that is the entire downside.
Defer and trigger, size fixed now. Do not rebuild the evaluation process this quarter. Pre-commit the trigger: the next time a decision made on this scorecard turns out badly, the weights get re-derived rather than adjusted. Decide now who runs that, because a review with no owner never happens.
Note the arrivals. The matrix and the ratio land today. Whether better weights produce better decisions cannot be known for as long as your outcomes take to arrive, which for investments is years, and that lag is precisely why nobody has audited these numbers.
What usually happens next
One further constraint is worth naming because it is the commonest reason this method is abandoned. Filling a pairwise matrix for five criteria means ten comparisons, and for seven it means twenty-one. That is genuinely tedious, and the tedium is the point at which people revert to asserting percentages. Budget the twenty minutes once rather than treating it as an ongoing cost, because the weights only need deriving when the criteria change, which is rarely.
Run the break test first. Has a rule changed, has an actor entered or left, has a measurement become a target? If your team now knows the weights, expect applicants and founders to optimise against them, and the criterion that was most predictive to become least so. Published weights stop measuring what they measured.
If nothing broke, the pattern is consistent. Scorecards drift toward whatever is easiest to score. Criteria that require judgement get compressed toward the middle of their range, and criteria with a hard number available spread across theirs, so the weighted total ends up driven by whichever inputs happen to be quantified rather than by the weights you set.
There is a second pattern worth expecting. Once the ratio fails and the comparisons are revised, people tend to revise toward whatever produces the weights they already wanted, and the second matrix passes. That is not the method failing, it is the method being used as a ratification device. The guard is to fill the matrix before looking at what weights it implies, which sounds trivial and is the entire discipline.
Subtract the counterfactual before crediting the scorecard. A process that has picked good companies in a rising market has not been tested. The question is whether it ranked them differently from a list sorted by the single most obvious metric, and usually it did not.
What this ensemble cannot see
All three lenses are about coherence. None of them is about correctness.
That is the load-bearing gap. A perfectly consistent set of weights derived from a flawless matrix can still be weighting the wrong criteria entirely, and the ratio will pass. This framework improves the internal logic of your judgement and says nothing about whether the judgement tracks reality, which is the thing you actually wanted.
There is also a real weakness in the central model, stated in its own card. Rank reversal was a serious critique and the debate about whether the results are arbitrary under hierarchical decomposition was never settled to everyone’s satisfaction. Ideal mode addresses the specific defect. It does not make the method uncontroversial, and anyone who rejected the whole approach on those grounds would have a defensible position.
And one property none of these models contains: the act of publishing weights changes the organisation. Once people know that team is weighted at thirty per cent, the meeting stops being about the company and starts being about the score, and the informal judgement that was quietly doing the real work stops being voiced.
The one action that survives the ignorance: before the next decision on this scorecard, fill the pairwise matrix yourself and compute the ratio. If it clears, you have learned your weights are at least coherent. If it does not, you have learned that a number currently deciding real outcomes has never been checked, and you learned it in half an hour.
Who has to move
The person who needs this owns the process, and they inherited the weights from someone who inherited them. The cheapest first test is one person, one matrix, half an hour, before any group meets. If the ratio fails, that is a far stronger argument for redesign than any objection to the criteria themselves.
Sources and notes
- Thomas L. Saaty, Thinking With Models: Mathematical Models in the Physical, Biological and Social Sciences. Chapter 8 develops hierarchies and priorities: paired comparisons on the one-to-nine ratio scale, the normalised principal eigenvector of the comparison matrix as the vector of weights, and the column-normalisation approximation that avoids computing the eigenvector directly. The consistency index, defined as the largest eigenvalue minus the number of criteria over that number minus one, is set out in the same chapter, along with the comparison against random-matrix averages to produce the consistency ratio and the guidance to revise where the ratio is considerably higher than ten per cent. The observation that one rarely needs to go beyond a seven by seven matrix in order to keep consistency relatively high is also section 8.3.
- Matthew A. Cronin, Cleotilde Gonzalez and John D. Sterman, Why don’t well-educated adults understand accumulation? A challenge to researchers, educators, and citizens, Organizational Behavior and Human Decision Processes 108(1), 2009, pages 116 to 130. Author copy: https://www.mit.edu/~jsterman/CroninGonzalezSterman061210.pdf. The abstract states that highly educated people are often unable to infer the behaviour of simple stock-flow systems, and that in a series of experiments persistent poor performance is not attributable to an inability to interpret graphs, contextual knowledge, motivation, or cognitive capacity.
A note on citing a contested method. The rank-reversal critique of this technique is serious and its author never fully conceded it. The method appears here anyway, in its repaired form and with the dispute stated in the body rather than buried, because the alternative currently in use is weights invented in a meeting and never tested at all. A contested procedure with a stated failure mode is a better instrument than an unexamined assertion.
Joshua Agonya Pi’Rwot, Founder.