Articles · 12

Practise AI decisions with feedback

Leaders practise presentations. They almost never practise the decisions that matter most.

The StepSim team, AIR APAC

Illustration: a leader in weekend clothes at a small cafe table on a shophouse walkway on a Saturday morning, laptop open, coffee beside her, a shuttered window behind.Illustrative scene

Think about what leaders rehearse. A speech to the board. A difficult conversation with a team member. A pitch to a new client. We practise these because we know the first attempt is rarely the best one.

Now think about what leaders do not rehearse. Deciding how far to trust an AI forecast before a budget meeting. Setting the limits for an AI agent. Questioning an AI summary that everyone else in the room has accepted. These decisions are new, high-stakes and growing in number. Most leaders meet them for the first time when they are real.

Why real work cannot teach this

You might expect people to learn good habits with AI on the job. Over time, they will see what works.

The problem is feedback. In most real work, you never find out whether the AI was right.

If you accepted the AI's supplier choice, you never see how the other supplier would have done. If you overruled its forecast, you never see what would have happened if you had followed it. When things go wrong, many causes are mixed together. When things go right, nobody asks why.

So people do not get clear feedback on their AI habits. They get noise. Habits form anyway, but they form without any signal about whether they are good.

A rehearsal knows the quality of the advice

A rehearsal can do what real work cannot. A business simulator is an online exercise where you take a work role, review information and make decisions. In a designed case, the designers know when the AI's advice is sound and when it is flawed.

That is different from one right business decision. A case can allow several reasonable decisions. What the case knows is the quality of the advice in front of you.

So after the rehearsal, a person can see things that real work hides. Which advice did I accept, and was it sound? What did I check, and what did I let through? What followed from the choices I made?

This idea is not new. John Sterman wrote about management flight simulators in his 2000 book Business Dynamics. They let managers try decisions that would be too costly to learn from for real. What is new is that the decision now includes an AI partner, and the case knows when the AI is wrong.

The debrief is where the value is

A rehearsal alone is not enough. The value comes from the debrief: a trainer-led discussion after the simulation, where people look back at their decisions together.

Scott Tannenbaum and Christopher Cerasoli reviewed the research on debriefs in 2013. Across 46 samples, debriefs improved performance by about 25%. That is general research on debriefs. It does not measure the effect of a StepSim session.

A good debrief does three things. It looks at the route before the result, so that luck does not decide the lesson. It compares different routes through the same case, so people see that there was more than one reasonable path. And it stays safe, so that people talk about what they actually did, not what they wish they had done.

What leaders can do

  • Pick the decisions to rehearse. Which AI decisions in your organisation are new, costly and rare? Those are the ones to practise.
  • Use a case where the quality of the advice is known. Without it, a rehearsal becomes a discussion of opinions.
  • Debrief the route, not the score. Ask what each person checked, accepted and challenged. Avoid ranking people.
  • Say up front who sees the record. People rehearse best when they know who will read their decisions, and why.

Pilots rehearse rare failures many times before they meet one. Leaders now face a new kind of decision with AI, many times a week. It makes sense to practise it somewhere safe first.

Sources

  • Sterman (2000). Business Dynamics: Systems Thinking and Modeling for a Complex World.
  • Tannenbaum and Cerasoli (2013). Human Factors 55(1):231–245.

Keep reading

More articles

Illustration: an experienced recruiter with short grey hair leaning back at her open laptop in the afternoon sun, a stack of untouched files beside it.Illustrative scene

03

When better AI makes experienced people worse

In a field experiment, recruiters given higher-quality AI were less accurate than those given weaker AI. The reason matters for every leader rolling out AI.

Put the question into practice.

The Friday file puts you in charge of an AI rollout decision, with evidence to review and a deadline to meet.