The AI adoption dashboard says 78 per cent weekly active users, up from 41 per cent. Licences are being consumed. The rollout is, by every number on the slide, a success.
Then someone asks what it bought, and the room goes quiet — because the honest answer is that nobody knows, and the metric that was chosen cannot answer the question. Weekly active users tells you a tool was opened. It does not tell you whether the work got better, worse, or simply got done in a different order.
This is a familiar failure with an unfamiliar twist. Usage data does not merely overstate AI adoption. It actively conceals the two findings that matter most.
Reason one: the average is the least informative number available
In one of the largest field studies of generative AI at work, Erik Brynjolfsson, Danielle Li and Lindsey Raymond examined the staggered rollout of an AI conversational assistant across 5,179 customer support agents. Access raised productivity, measured as issues resolved per hour, by 14 per cent on average.
That average is close to meaningless on its own. The same study found a 34 per cent improvement among novice and low-skilled workers and minimal impact on experienced, highly skilled ones. The mechanism the authors describe is that the tool disseminates the best practices of the more able workers and helps newer staff move down the experience curve.
So “14 per cent” is not a description of what happened to anybody. It is an artefact of the mix. Two organisations with identical headline gains can have completely different underlying stories — one where new joiners reached competence in half the usual time, another where a few experienced people got marginally faster and everyone else changed nothing.
Reporting the mean is the single most common measurement error in AI rollouts, and it destroys the most valuable finding in the data.
Reason two: the same usage contains gains and damage
This is the finding that should change how you measure, and it comes from a preregistered randomised experiment with 758 management consultants at Boston Consulting Group, reported by Fabrizio Dell’Acqua and colleagues as Navigating the Jagged Technological Frontier.
On eighteen realistic tasks inside the frontier of the model’s capability, consultants using AI completed 12.2 per cent more tasks, 25.1 per cent more quickly, at significantly higher quality. On a complex managerial task deliberately chosen to sit outside that frontier, consultants using AI were 19 per cent less likely to produce correct solutions than those working without it.
Same people. Same tool. Same week. Opposite outcomes, determined entirely by which side of an invisible and irregular boundary the task fell on — which is what the authors mean by a jagged frontier.
Now consider what a usage metric does with that. Both groups show as active users. Both consume licences. Both appear in the adoption percentage as successes. The metric is structurally incapable of separating the 25 per cent faster from the 19 per cent more wrong, and an organisation optimising for higher usage is pushing on both at once.
Reason three: the denominator is wrong
Sanctioned-tool telemetry misses everything happening outside the sanctioned tool, and that is not a rounding error. Microsoft and LinkedIn’s 2024 Work Trend Index, surveying 31,000 knowledge workers across 31 countries, found 78 per cent of AI users bringing their own tools to work, and 52 per cent reluctant to admit using AI for their most important tasks.
Your platform metrics therefore measure a self-selected subset of AI use, biased away from exactly the high-stakes work you most need to understand. Shadow AI is not a separate governance topic; it is a hole in your measurement.
| The metric | What it is taken to mean | What it actually establishes |
|---|---|---|
| Weekly active users | People have adopted it | A session was opened. Nothing about the task, the output, or whether it was used |
| Prompts per user | Depth of engagement | Volume — which also rises when people are struggling to get a usable answer |
| Licences consumed | Value delivered | Procurement worked |
| Time saved (self-reported) | Productivity gain | What people believe, from a population half of whom conceal their most important use |
| Satisfaction score | It is working | People enjoy using it. Enjoyment and accuracy are unrelated |
Measure the frontier deliberately
If performance depends on which side of the frontier a task sits, then the frontier is the thing to map — and no vendor can map it for you, because it is specific to your work, your data and your models.
The exercise is smaller than it sounds. Take one function. List the ten to fifteen tasks that consume most of its time. For each, run a short structured comparison: the same task done with and without the tool, scored for correctness by someone competent to judge, not for speed alone. You are producing three groups:
- Inside the frontier — faster and at least as good. Push adoption hard here; this is where the return is.
- Outside the frontier — slower, or faster and wrong. Say so explicitly, in writing, to the people doing the task. This is the list that prevents the 19 per cent.
- Ambiguous — good output that requires expert checking to validate. Usable, but only with a checking step defined, and worthless where the checker is the bottleneck.
That three-way list is worth more than any dashboard, and it is the artefact most AI programmes never produce. It also has to be revisited when models change, because the frontier moves — which is an argument for keeping the exercise small enough to repeat.
Segment by prior skill, always
Given the Brynjolfsson result, reporting AI impact without splitting by experience level is close to negligent. The split is not noise to be averaged away; it is the primary finding, and it changes what you should do next.
If your gains are concentrated among newer staff, the value case is about time-to-competence, onboarding and reduced dependence on senior colleagues — a capability story that also relieves the load on the small number of experts everyone routes questions to. If your gains are concentrated among experienced staff, you are probably automating a bottleneck, and the benefit will be realised only if the freed capacity is redeployed rather than absorbed.
Those are different investments with different follow-on actions, and a single headline percentage tells you which one you are in exactly never.
Measure quality, or you are measuring nothing
Throughput without a quality measure is the most dangerous number in an AI programme, because AI reliably improves throughput and only sometimes improves quality. A dashboard showing volume alone will look excellent during the period in which error rates are quietly rising.
Three practical quality measures, in ascending order of effort:
- Rework rate. How often does AI-assisted output come back? Usually already captured somewhere — amendments, credit notes, resubmissions, corrections.
- The planted-error test. Give people a realistic AI-generated output for their own role containing one obvious and one plausible defect, and record the proportion who catch the plausible one. Twenty minutes, comparable across functions and repeatable over time.
- Blind sampling. A reviewer scores a sample of output without knowing whether AI was involved. The only method that produces a defensible quality number, and worth the cost for anything customer-facing or regulated.
The middle one is the best value by a wide margin, and it doubles as a genuine test of AI literacy rather than of attendance.
Do not let the metric become the goal
Once weekly active users becomes the reported measure of an AI programme’s success, it will rise. Teams will be encouraged to use the tool, usage targets will appear in objectives, and the number will detach from the thing it was standing in for.
This is the ordinary fate of a load-bearing indicator, and the defence is the same one that applies to any readiness criterion: never carry a construct on a single measure. “AI is creating value in this function” should be evidenced by a task-level frontier map, a segmented productivity measure and a quality measure that disagree with each other when something is wrong. One number cannot be interrogated. Three can.
A defensible AI adoption measure
Putting it together, what a credible report contains:
- Which tasks the tool is being used for, from the frontier map — not how many sessions were opened.
- Effect on those tasks, split by prior skill level, with speed and quality reported together and never separately.
- A named estimate of what is happening outside the platform, because a measure that ignores unsanctioned use is describing a minority of the behaviour.
- The tasks explicitly ruled out, and evidence that the people doing them know.
- What the freed capacity was used for. Time saved is not a benefit until something else happened with it — the distinction that decides whether adoption actually caused the benefit.
That report is harder to produce than a usage chart and considerably harder to make look good, which is the main reason it is rare.
But an organisation that can say our claims-handling drafting is 25 per cent faster with no rise in rework, the gain is concentrated in staff with under two years’ service, and we have ruled out three task types where accuracy fell understands its own adoption. An organisation reporting 78 per cent weekly active users knows how many people opened a tab.
More on measuring adoption honestly in the AI Adoption Readiness Hub, or use the AI Adoption Readiness Assessment to establish a baseline before the usage dashboard sets the narrative.
