FounderWiseDecisions, not feeds
← All articles FounderWise · Long-form

Less than 20%, and the questions the model got wrong

A Census Bureau story from 26 May 2026 puts firms with four or fewer employees under 20% AI use. Two experiments show what happens on the questions the model got wrong.

07 Oct 2026 5 min read By Joshua Pi’Rwot
Share X LinkedIn
Less than 20%, and the questions the model got wrong

Less than 20% of US firms with four or fewer employees reported using AI. That line is in a Census Bureau story dated 26 May 2026, written by Adam Grundy, Cory Breaux and Dhanapati Khatiwoda. The survey is the Business Trends and Outlook Survey. The authors reviewed collections from 14 December 2025 through 3 May 2026.

The story prints a ceiling, not a point. I am not going to call it 19. The share for that band sits somewhere under the line they published.

What moved, and what held still

From December 2025 to May 2026, use increased among firms with at least 20 employees. It did not change significantly among firms with fewer than 20.

Firms with at least 250 employees: 37% reported using AI. That 37% is not pinned to the single collection ending 3 May. For the period ending 3 May 2026, 32% of firms with 100 to 249 employees reported using AI. The national rate that period was 19.8%. Information was at 39.7%. Finance and Insurance was at 33.9%. Neither sector had shifted significantly since December. Retail trade sat around 14%, and about 17% expected to use AI in the next six months.

As of 3 May 2026: Information 39.7%, Finance and Insurance 33.9%, national rate 19.8%, retail trade around 14%.
Same story, 3 May 2026. Around 14% is the story’s phrase, not a point.

Across all firms, over those six months, use hovered between 17% and 20%. Between 20% and 23% expected to use it in the next six months.

On 17 November 2025 the question changed. Before, it asked about AI use in producing goods or services. After, it asked about use in any business function. The line about four-person firms answers that wider question.

Small firms against larger firms. Four or fewer: less than 20% reported using AI. 250 or more: 37%.
Census Bureau story, 26 May 2026. Less than 20% is a ceiling. The 37% for firms with 250 or more employees is not pinned to the collection ending 3 May. The 32% is.

The story cannot see one shop’s books. For firms under 20 employees, it says the rate did not change significantly across those six months.

The questions the model got wrong

Once a firm does use a model, one version of the advice puts the reasoning and the sources on the screen. Another adds a mark for how sure the model is. The two studies report the questions the system got wrong, and the questions it got right. On the questions it got right, the tool helped.

Zana Buçinca, Maja Barbara Malaya and Krzysztof Z. Gajos published the first study in April 2021. People looked at a plate of food, picked the ingredient highest in carbohydrates, and replaced it with something that cut the carbs and stayed close in flavor. The model was a simulation, set at 75% accuracy. When it was wrong, it missed the high-carb ingredient, which the experimenters had made obvious, and pointed at a low-carb one instead. After the filters, 199 people remained.

On all questions together, both assisted groups beat the group with no model. Overall performance was 0.17 with no AI, 0.35 with a simple explanation, and 0.33 when people had to do something before accepting the suggestion.

On the incorrect predictions, the order reversed. For which ingredient held the carbs: no AI 0.49, simple explanation 0.08, forcing 0.27. Standard errors 0.04, 0.03 and 0.03. F(2, 242.6) = 35.59. Forcing means the designs that made a person act before simply reading the suggestion. It beat the plain explanation. It did not catch the unaided group.

On the carb-source questions the model got wrong: no AI 0.49, forcing 0.27, simple explainable AI 0.08.
Carb-source questions the model got wrong. Buçinca, Malaya and Gajos, 2021. This figure is the misses only.

The second study is Sunnie S. Y. Kim, Q. Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard and Jennifer Wortman Vaughan, presented in Rio de Janeiro from 3 to 6 June 2024. The final sample is 404 people, and eight medical questions. The system was right on four and wrong on four, by design. That split is not a measured accuracy rate out in the world.

Overall, people with no system scored 74.2%. People with an unhedged system scored 63.9%. People who saw a first-person hedge scored 72.8%. One of the sentences in the materials is “I’m not sure, but it seems like.” A general-voice hedge landed at 67.9%. The contrasts the authors mark are the unhedged system below the first-person group, and the unhedged system below no system. I am not claiming a separate significance mark for the general voice.

Overall percent correct: no system 74.2, first person 72.8, general voice 67.9, control 63.9.
One score for the eight questions. The marked tests are control below first person, and control below no system.

When the system was right, seeing it helped: 88.5%, against 77.9% with no system. When it was wrong, seeing it hurt: 33.0%, against 64.7% with no system. The hedges narrowed the hole on the wrong questions. The authors say both hedge groups still finished those questions behind the people who had no system.

When the system was right: control 88.5%, no system 77.9%. When it was wrong: control 33.0%, no system 64.7%.
Four of the eight were wrong by design. Seeing the system helped on the hits and hurt on the misses.

What the two pages will and will not support

Both samples are US crowdworkers. One task is a plate of food, with the high-carb ingredient made obvious. The other is eight medical questions with the errors assigned in advance. The medical sample, the authors say, is younger and more educated than the country. None of that is a ledger, a tax filing, or a supplier quote.

If the firm has fewer than 20 people, the Census story places it in the band whose reported use did not move from December 2025 to May 2026. At four employees or fewer, the band is the one the Bureau put under 20%.

If a model is in use on a question that has a wrong answer, write yours down before you open the output. On the plates the model got wrong, that step scored 0.27. Reading the explanation and going with it scored 0.08. No model scored 0.49.

Sources

The story’s data page and working paper were not opened. The paper Kim cites on a first-person sentence amplifying a confident claim was not re-read. No number in this piece comes from those.

Lock in your calls.

You’ve marked 0 of 5. Now choose how often you want the signals.

Step 1 · Pick your cadence

The DispatchWeekly · your Monday 5 callsFreealways

Step 2 · Where to send it

Personalize your BriefThe Brief

Tune every edition to the markets and industries you actually act on.

🔒 Unlock personalization — The Brief, $19.99/mo →
Free Dispatch forever · upgrade anytime · we never share your details.
Know where you stand?
The Dispatch tells you what changed. Knowing what to do about it is a different question, and it is the one FounderWise answers. Start with the free Traction Audit: 12 questions, about 3 minutes, scored out of 100.
Find out where you stand →

For teams, syndicates & programs

Recommended
Team
$15/seat · mo
Daily Brief for the whole team (min 3 seats).
  • Everyone on the same signal
  • Admin + shared watch-list
  • One invoice · ~25% off solo
Get Team →
Cohort Licence
$2,900/yr
Co-branded seats for one cohort, for accelerators, funds & programmes.
  • Up to 25 founder seats
  • Your logo, your cohort
  • The record of what your cohort committed to
Talk to us →
Pass the Dispatch on
Know a founder making these calls blind? Send them this week’s five — free, every Monday.

Decisions, not feeds. · Curated by Joshua Pi’Rwot · FounderWise · Free Audit · Store · parent of Business Growth Accelerator

Call committed. We’ll hold you to it.
Know where you stand →