The pilot was a clear success. Twelve people in one team, six weeks, a 40 per cent reduction in drafting time and enthusiastic feedback. The business case was approved on the strength of it.

Nine months later the same capability is available to 1,400 people, adoption has plateaued somewhere around a third, and nobody can find the 40 per cent anywhere in the operating numbers.

Nothing was faked. The pilot result was real. It was simply not a prediction of the rollout, because a pilot and a rollout differ in at least five ways that all point in the same direction — and none of them are usually written down.

First, about that statistic

You will have seen the claim that 95 per cent of generative AI pilots fail. It comes from a 2025 MIT-affiliated report, and it has been repeated in headlines everywhere.

It is not usable as evidence. The figure rests on 52 executive interviews, a survey of 153 leaders and a review of around 300 public deployments, and it has attracted substantial criticism on exactly the points that would matter: how failure was defined, whether efficiency and cost effects were counted at all, the transparency of the underlying data, and a gap between what the source document says and how it was reported.

The underlying phenomenon is real and widely experienced. The number is not a measurement of it. That distinction matters here more than usual, because a headline figure of that size encourages the conclusion that AI does not work — when the better-evidenced conclusion is that pilots and rollouts are different problems and are routinely treated as the same one.

Mechanism one: pilots select tasks inside the frontier

This is the largest effect and the least discussed.

In a preregistered randomised experiment with 758 management consultants at Boston Consulting Group, Fabrizio Dell’Acqua and colleagues found that on eighteen tasks inside the model’s capability frontier, AI users completed 12.2 per cent more tasks, 25.1 per cent faster, at higher quality. On a complex managerial task chosen to sit outside that frontier, AI users were 19 per cent less likely to produce a correct solution.

Now consider how a pilot is scoped. Someone picks a use case they are fairly confident will work. That is sensible project management and it is also, unavoidably, selection on the outcome: pilots are drawn overwhelmingly from inside the frontier.

A rollout is not scoped that way. It hands the tool to everyone and lets them apply it to whatever they are doing, which is a mixture of inside-frontier tasks, outside-frontier tasks and ambiguous ones. The average effect of the rollout is therefore mathematically certain to be worse than the pilot, even if nothing at all went wrong operationally.

Mechanism two: pilots are staffed by volunteers

Pilot participants put their hands up. They are curious, tolerant of rough edges, willing to try a second prompt when the first fails, and predisposed to want the thing to work.

The rollout population includes the person who was on leave during the briefing, the person who has been doing this job successfully for nineteen years and does not see the problem, and the person who tried it once, caught it inventing something, and quietly stopped.

That last group is a documented failure mode, not an attitude problem. Human factors research on automation calls it disuse — the neglect or underuse of a system, classically triggered by an early false alarm. It is invisible in adoption dashboards, which count active users rather than the people who actively stopped.

Mechanism three: the gain accrues to people who were not in the pilot

Brynjolfsson, Li and Raymond’s study of 5,179 customer support agents found a 14 per cent average productivity gain that decomposed into a 34 per cent improvement for novice and low-skilled workers and minimal impact on experienced, highly skilled ones.

Pilots are usually run with capable, senior, engaged people — precisely the population where that research found the least benefit. Two things follow, and they pull in opposite directions:

You cannot know which of these you have unless the pilot deliberately includes both populations and reports them separately. Almost none do, which is why reporting a single average is the most common measurement error in AI programmes.

From pilot to rollout Pilot went well and the rollout is stalling? Book a 20-minute scoping call to work out which of the five gaps you are actually facing, and what closes it.
Book a 20-minute scoping call

Mechanisms four and five: support and measurement

Support asymmetry. A twelve-person pilot has an attentive human somewhere in the loop — a product owner in the channel, a vendor consultant on a weekly call, someone who notices when a participant goes quiet. At 1,400 users that person does not exist, and the ratio problem is severe: the thing that made the pilot work is the first thing that does not scale.

Measurement change. Pilots are evaluated on enthusiasm, self-reported time saving and anecdote, all of which are legitimate for a pilot and none of which survive contact with an operating review. When the rollout is judged on cost per case or output per head, it is being asked a question the pilot never answered.

 PilotRollout
TasksChosen because they were likely to workWhatever people happen to be doing
PeopleVolunteers, engaged, tolerant of frictionEveryone, including the sceptical and the absent
SupportAttentive, roughly 1:12A mailbox, roughly 1:1,400
FailureInteresting; feeds the next iterationCostly, visible, and attributed to users
Measured byEnthusiasm and self-reported time savedCost per case, output per head, error rates

Read down the two columns and the honest conclusion is that a successful pilot is weak evidence for a rollout. It establishes that the technology can work somewhere, for someone, under favourable conditions. That is genuinely useful and it is not what it gets used for.

Design the pilot to answer the rollout’s questions

The fix is not a bigger pilot. It is a pilot that produces different outputs.

A pilot run this way often looks worse than a conventional one. It should. It is producing an estimate rather than a demonstration.

The rollout is a change programme, not a distribution exercise

The deeper problem is categorical. Because the pilot was a technology exercise that succeeded, the rollout is planned as the same technology exercise at larger scale — provisioning, licences, an enablement session, a launch communication.

But everything that differs between the two columns above is a people and process problem. Task suitability is process design. Support ratios are capacity planning. Sceptical users are a capability and trust question. Measurement change is governance.

Microsoft and LinkedIn’s 2024 Work Trend Index found 79 per cent of leaders agreeing that AI adoption is necessary to stay competitive, while 60 per cent said their company lacked a plan and vision to implement it. That gap is not a technology gap. It is what happens when a successful demonstration is mistaken for an implementation plan — and it is why AI needs strategic change management rather than enablement alone.

What a pilot should hand over

A pilot that ends with “it worked, let’s scale” has produced a decision and no information. A pilot that has done its job hands over five things:

None of that requires more time than a conventional pilot. It requires deciding at the outset that the pilot exists to produce an estimate rather than a business case — which is an uncomfortable reframing, because the pilot is usually being run by people who need it to succeed.

That, rather than any headline failure rate, is the honest explanation for the gap. Pilots succeed because they are designed to. Rollouts fail because nobody designed them at all.

More on leading AI-enabled change in the AI Adoption Readiness Hub, or use the AI Adoption Readiness Assessment to size the rollout conditions before the pilot result sets expectations.

Ritvars Mētra

Ritvars Mētra

Founder of ReadinessCompass

Ritvars Mētra is the founder of ReadinessCompass, where he develops practical tools for understanding and managing organisational change complexity. His work focuses on adoption readiness, stakeholder analysis, and evidence-based change management for large-scale software and AI implementations.

View full profile →