FounderWiseDecisions, not feeds
← All articles FounderWise · Long-form

They were looking straight at the evidence that said no

Of 18 people who followed a wrong automated diagnosis, 11 had checked every parameter, including the ones that said no. Showing people more evidence is not the fix.

04 Oct 2026 5 min read By Joshua Pi’Rwot
Share X LinkedIn
They were looking straight at the evidence that said no

There is a study I keep coming back to, because it takes away the answer most of us reach for.

Participants ran a simulated process plant with an automated aid that diagnosed faults for them. At one point the aid gave a confident diagnosis that was wrong. Eighteen people followed it.

The obvious explanation is that they got lazy and skipped the checks. That is the explanation I would have given, and it is the one every article about this writes. The researchers had the logs, so they went and looked.

Eleven of the eighteen had completed the verification in full. They had opened every parameter they were supposed to open, including the ones that contradicted the diagnosis, and they followed the aid anyway.

Of 18 people who followed a wrong automated diagnosis, 11 had completed the verification in full

What the numbers leave

Seven of the eighteen had not finished checking. Of those seven, only two had never opened the screen that falsified the aid’s answer.

So of eighteen wrong decisions, sixteen were made by people who had the disconfirming evidence available to them. Two out of the whole supported group of forty two were genuinely unable to know.

Meanwhile the control group, running the same fault with no automated aid at all, got it right thirteen times out of fourteen.

Juliane Reichenbach, whose dissertation at TU Berlin carries these numbers, does not call it discounting. She calls it a looking but not seeing effect, and draws the comparison to inattentional blindness. Information that is looked at is not processed. The eye passes over the number that contradicts the machine and the brain files it as agreement.

Why this ruins the easy fix

Almost every proposal for safer AI use is a variant of one idea: show the person more. Show the confidence score. Show the sources. Show the reasoning. Make the evidence available and the human will catch the error.

This study says availability was never the constraint. The evidence was available to sixteen of the eighteen people who got it wrong.

There is a second study, by researchers at Harvard, that tests the stronger version of that idea directly. Participants made decisions with an AI that was sometimes wrong. On the cases where the AI was wrong, the group with a plain explainable AI interface got 8 percent right. The group with a cognitive forcing design, which made them commit to an answer before seeing the AI’s, got 27 percent. The group with no AI at all got 49 percent.

Adding the explanation did not help. It made people worse than having no assistance, by a wide margin. Forcing a commitment recovered part of it and did not get anywhere near the unaided group.

What actually moved the number

In the same body of work, one intervention worked properly.

Participants who had previously watched the aid fail made the error at 4.5 percent. Participants meeting their first false diagnosis made it at 20.5 percent. Having seen the machine be wrong once cut the error rate by about three quarters.

Not an explanation. Not a confidence score. An experience of the thing being wrong.

What this means if you run your tools alone

If you have no one checking behind you, three practical things follow.

A check you can pass while agreeing is not a check. Reading the output and nodding is exactly what the eleven did, and they were doing it properly by every procedural standard. If your verification step can be completed without ever producing a disagreement, it cannot tell you anything.

Write your answer before you look at the tool’s. That is the only intervention in this literature that reliably recovers anything, and it costs nothing. Decide, then open the output, then compare. The comparison is the check. Reading the output and then forming a view is not.

Break your own tool on purpose, once. Feed it something you know it gets wrong and watch it be confidently wrong. The single biggest measured improvement in this whole literature came from having seen the failure, not from being warned about it. A team that has never watched its tool fail has a team-wide 20 percent problem.

The uncomfortable part

I have written before that a system can report success while doing nothing. That one is solvable: build a better witness, ask an independent source, check the thing that actually certifies the outcome.

This is harder, because here the witness worked. The contradicting number was on the screen. The verification was complete. The person looked directly at the evidence that said no, and said yes.

You cannot fix that by adding another display. You fix it by changing the order of operations, so your own judgment is on the table before the machine’s arrives, and by having watched the thing be wrong at least once so that possibility is live in your head rather than theoretical.

Before you hand another check to a tool, the AI Readiness Check scores whether that process is ready for it. Twelve questions, four minutes, no email required: founderwise.io/ai-readiness/

If you want to go through the checks in your own business with someone who works on this, book a strategy call: cal.com/pirwot/strategy-call

Advice is free. A verification step that cannot disagree with you is worse than free, because it costs you the belief that you checked.

So: the last time one of your tools was wrong, did you catch it, or did someone else?

Sources

Lock in your calls.

You’ve marked 0 of 5. Now choose how often you want the signals.

Step 1 · Pick your cadence

The DispatchWeekly · your Monday 5 callsFreealways

Step 2 · Where to send it

Personalize your BriefThe Brief

Tune every edition to the markets and industries you actually act on.

🔒 Unlock personalization — The Brief, $19.99/mo →
Free Dispatch forever · upgrade anytime · we never share your details.
Know where you stand?
The Dispatch tells you what changed. Knowing what to do about it is a different question, and it is the one FounderWise answers. Start with the free Traction Audit: 12 questions, about 3 minutes, scored out of 100.
Find out where you stand →

For teams, syndicates & programs

Recommended
Team
$15/seat · mo
Daily Brief for the whole team (min 3 seats).
  • Everyone on the same signal
  • Admin + shared watch-list
  • One invoice · ~25% off solo
Get Team →
Cohort Licence
$2,900/yr
Co-branded seats for one cohort, for accelerators, funds & programmes.
  • Up to 25 founder seats
  • Your logo, your cohort
  • The record of what your cohort committed to
Talk to us →
Pass the Dispatch on
Know a founder making these calls blind? Send them this week’s five — free, every Monday.

Decisions, not feeds. · Curated by Joshua Pi’Rwot · FounderWise · Free Audit · Store · parent of Business Growth Accelerator

Call committed. We’ll hold you to it.
Know where you stand →