An organisation announces an AI programme and use cases multiply overnight. Every function wants a place on the list. The working group tries to maintain enthusiasm, so it approves ten pilots, then fifteen. Three months later the programme has produced activity, but no one can say which result deserves a decision.

This is how pilot portfolios become a form of avoidance. A large number postpones the uncomfortable act of choosing. It spreads attention across too many teams and gives weak ideas somewhere to hide.

A pilot should reduce uncertainty around a decision. If no decision follows, you ran a demonstration.

Start with a real problem

‘We want to use AI in recruitment’ is not a pilot. It names a technology and a function. A useful brief names a person, a recurring piece of work and a measurable difficulty. It makes clear what evidence would justify changing the way the work is done.

The strongest use cases often sound modest. They shorten the preparation for a recurring decision, improve the consistency of a review or help a team examine more options before a human makes the call. Their value comes from proximity to work, not theatrical novelty.

Write the controls before the excitement arrives

Every pilot needs a named data class and a human reviewer who can recognise a bad answer before it reaches a colleague or customer. These details belong in the original brief. Adding them after the prototype works turns governance into a last-minute obstacle and encourages teams to treat safety as somebody else’s delay.

A simple pilot brief should answer the following questions:

  • What question are we testing?
  • Who experiences the problem now?
  • What information will the pilot touch?
  • Who reviews the output and remains accountable?
  • What result would justify continuing?
  • What result would make us stop?

Protect the manager who owns the test

Pilots fail quietly when managers are asked to run them on top of unchanged delivery targets. The responsible manager must say what the pilot displaces and protect time for review. Without that decision, the team learns that experimentation is encouraged only after the real work is finished.

The manager also needs permission to stop. A kill criterion is useful only when ending a weak pilot is treated as good management rather than a loss of face. The purpose of the test is learning. Sometimes the most valuable result is evidence that the idea should go no further.

Three creates comparison

One pilot can become a story about a charismatic team. Three let you compare conditions. You can see whether a result travels across functions, whether one data class creates different friction and whether manager behaviour changes the outcome.

More can come later. The opening set should be small enough for the working group to stay close to the work and learn why each result occurred. Choose three problems worth understanding. Write the decision each one must inform. Then give the teams enough attention to produce evidence you can trust.

This essay draws on Marc’s forthcoming field guide, The Missing Piece: How HR and L&D Turn AI Strategy into Everyday Practice.