Articles · 03

When better AI makes experienced people worse

It sounds wrong, but the research is clear. A better tool can make good people stop checking.

The StepSim team, AIR APAC

Illustration: an experienced recruiter with short grey hair leaning back at her open laptop in the afternoon sun, a stack of untouched files beside it.Illustrative scene

Most AI rollouts rest on a simple belief. If the tool gets better, the work gets better. Buy the stronger model, and quality will follow.

One field experiment tested that belief, and it did not hold.

The recruiters

Fabrizio Dell'Acqua ran a pre-registered field experiment with 181 professional recruiters. Each recruiter reviewed CVs with help from an AI. Some were given higher-quality AI. Some were given lower-quality AI.

You might expect the first group to do better. They did not. The recruiters given higher-quality AI were less accurate. They also put in less effort than the recruiters given lower-quality AI.

The working paper has a memorable title: Falling Asleep at the Wheel. The image fits. When the AI was good most of the time, people relaxed. They stopped reading closely. So when the AI was wrong, fewer people caught it.

With weaker AI, people stayed alert. They knew the tool made mistakes, so they kept checking. Their own effort made up for the tool's limits.

Experts are not protected

It is tempting to think this only happens to junior staff. It does not.

A study by Rosbach and colleagues looked at 28 trained pathology experts. These were pathologists, pathology residents and pathology staff. They made estimates with help from an AI. In about 7% of the AI-assisted estimates, 38 of 560, they showed automation bias. In plain words, they went along with AI advice that was wrong.

Two other details matter. First, AI improved their results overall. The tool was useful. Second, time pressure did not change how often the bias happened. But it appeared to make each case more severe.

So the lesson is not "AI is bad". The lesson is that a useful tool and careless reliance arrive together, even for experts.

Why this matters now

The AI your people use today is much better than the AI of two years ago. It writes clearly. It sounds sure. It is right a large share of the time.

That is exactly the condition the recruiters faced. The better the tool, the easier it becomes to stop checking. The errors that remain are fewer, but they are harder to see, because they sit inside work that looks finished.

This is one way people fall behind their tools. Their skill does not disappear overnight. It goes quiet, because the tool rarely needs it. Then, on the day it is needed, the habit of checking is not there.

What leaders can do

A leader cannot stop the AI from improving, and should not want to. The task is to keep human attention where it matters.

  • Say where the AI is weak. Every tool has areas where it is less reliable. Name them for your team, in plain words.
  • Keep some checks deliberate. For high-stakes work, ask for one specific check against a source, every time, even when the AI is usually right.
  • Ask about effort, not only output. In reviews, ask what the person checked. A good paper with no checks behind it is a risk, not a success.
  • Watch the quiet months. The risk is highest after a long run of good AI answers. That is when people relax.

Sources

  • Dell'Acqua, F. Falling Asleep at the Wheel. Working paper. PDF.
  • Rosbach, Ganz, Ammeling, Riener and Aubreville (2025). BVM 2025. arXiv.

Keep reading

More articles

Illustration: a leader in a white shirt at a whiteboard drawing empty boxes joined by arrows, two colleagues at the table listening.Illustrative scene

07

What does judgment with AI mean?

Judgment with AI means deciding what to use, what to check and what to do next. Five everyday examples, what StepSim makes visible, and where the idea comes from.

Put the question into practice.

The Friday file puts you in charge of an AI rollout decision, with evidence to review and a deadline to meet.