Two weeks after go-live somebody asks how it is going, and the answer is a usage figure. Eighty-three per cent of licensed users logged in last week.
That number is true, and it answers almost nothing. It does not say whether the work being done in there is correct, whether the organisation has settled, or whether next month will be calmer than this one.
Post go-live measurement has three distinct jobs — adoption, proficiency and stabilisation — and they run on different clocks. Blending them into one health score is the fastest way to lose the diagnosis.
Three constructs, not one
The separation is not a stylistic preference. The canonical framework for measuring whether an information system has succeeded — DeLone and McLean’s ten-year update to their IS success model, in the Journal of Management Information Systems — keeps use, quality and net benefits as separate but interrelated dimensions. Use was never the measure of success; it is one dimension among several, and a high value on it says nothing about the others.
Prosci’s framing of the people side splits the same way: speed of adoption, ultimate utilization and proficiency are three separate factors that each constrain return independently.
| Question | Measures | Readable from | |
|---|---|---|---|
| Adoption | Are the right people using it? | Usage by role, frequency, share of transactions in-system | Week 1 |
| Proficiency | Are they using it well? | First-time-right, rework, exception rate, unaided completion, time on task | Needs repetitions, not weeks |
| Stabilisation | Has it settled? | Variability of incident arrivals, novel-issue share, backlog age, workaround closure | After at least one full cycle |
The combinations matter more than the individual scores. High adoption with low proficiency is the most dangerous state on this list, because everyone is now putting poor-quality work into the system of record at volume — and it looks like success on a usage dashboard.
Adoption: breadth, weighted by who matters
This is the best-understood of the three, so keep it short.
Two refinements are worth making. Measure by role rather than by population, because an enterprise average hides the roles the benefit actually depends on. And weight by criticality rather than headcount — twelve people who maintain the master data matter more to outcomes than two hundred occasional viewers.
The characteristic failure is the compliance plateau: usage climbs quickly to whatever avoids attention and then stops. A curve that flattens early and low is not adoption succeeding slowly; it is people finding the minimum. The fuller treatment of what leaks in that gap is in why benefits leak after go-live.
Proficiency: measure against a trajectory, not a date
Here is the mistake almost every programme makes: setting a proficiency target against the calendar. First-time-right should reach 90% by week six.
Proficiency does not improve with time. It improves with repetitions.
Theodore Wright established the underlying relationship in 1936 in Factors Affecting the Cost of Airplanes, observing that each doubling of cumulative output reduced the labour required by a consistent percentage — the learning curve, later generalised as Wright’s law. We learn by doing, at a rate tied to how much doing has happened.
Two consequences for post go-live measurement, and both are practical:
- Plot proficiency against cumulative repetitions, not weeks. A clerk processing forty invoices a day and a controller running a task twice a month are on completely different curves. Judging both at week six tells you which process is high-volume, not which team is struggling.
- The signal is deviation from trajectory, not the absolute level. A monthly process still at 60% first-time-right after two cycles is entirely normal. The same figure after eight cycles is a finding. Flatlining is the alarm; a low starting point is not.
This also gives a defensible answer to the periodic-process problem. When a task runs monthly, people have had three attempts by the end of hypercare. Expecting fluency is not optimism, it is arithmetic error — and the correct response is engineered repetition (rehearsal on real cases) rather than more explanation, which is the same reason practice beats training.
Useful proficiency measures: first-time-right rate, rework volume, exception and rejection rates, correction transactions, time on task, and unaided completion sampled by observation rather than asked about.
Stabilisation: the one nobody defines
Every programme says it is waiting for things to stabilise. Very few can say what that means, which is why hypercare so often ends on a budget date instead of a criterion.
There is a rigorous definition available, and it comes from process control rather than change management. Walter Shewhart’s distinction between common-cause and special-cause variation, central to W. Edwards Deming’s work on knowledge of variation, holds that a process is stable when the variation remaining in it is common-cause only — the ordinary noise of the process — with no identifiable special causes still firing.
Translated: stabilisation is not a low number of incidents. It is a predictable one. An operation running at forty tickets a week, every week, within a narrow band, is stable and can be planned around. One averaging fifteen but ranging from three to sixty is not stable, even though the average looks better.
That gives four measures worth more than incident volume:
- Variability of arrivals, not just the mean. Week-to-week spread narrowing is the clearest stabilisation signal available.
- Novel-issue share. What proportion of this week’s incidents are a kind of problem not seen before? Early on it is most of them. Stabilisation is when new kinds stop appearing, even if volume is still moderate. This is the single most informative stabilisation measure and it costs one extra tag on the ticket.
- Backlog age. The oldest unresolved blocking issue. A stable operation does not accumulate an ageing tail.
- Workaround closure rate. Open workarounds falling means the organisation is converging on one way of working; flat means it is running two operating models and paying for both.
And one caution that matters more than any of them. Falling incident volume can mean stabilisation, or it can mean abandonment. People stop raising tickets when they have given up and built a workaround. The two look identical on an incident chart and are distinguished only by adoption and workaround data — which is why stabilisation cannot be read on its own.
Different clocks
Because the three become readable at different points, reporting all of them from day one produces two months of meaningless numbers and trains everyone to ignore the pack.
| Period | What is meaningful | What is not yet |
|---|---|---|
| Days 0–14 | Access and login, transaction volumes flowing, critical workflows completing end to end, incident arrivals and novelty | Proficiency — too few repetitions. Stabilisation — nothing has settled by definition |
| Weeks 2–8 | Adoption by role, first-time-right against repetitions, unaided completion, workaround register, novel-issue share falling | Business outcomes — still absorbing the transition dip |
| Weeks 8–16+ | Stabilisation across a full period-end, proficiency trajectory, workaround closure, first business-outcome movement | Full benefit realisation — that is a longer horizon still |
The right-hand column is the useful discipline. Saying explicitly that proficiency is not yet readable prevents somebody drawing conclusions from three data points, and prevents the opposite failure of quietly reporting it as green because nothing has gone visibly wrong.
Note also that at least one full cycle — a month-end, a quarter-end, a payroll run — must complete before stabilisation means anything. A system that has never been through period-end has not been tested at the point it is most likely to fail.
Four ways this goes wrong
- One blended score. Adoption, proficiency and stabilisation averaged into “post go-live health: amber.” The three need different owners and different interventions; the blend tells nobody which to run.
- Declaring victory on usage. High adoption with unmeasured proficiency is how bad data enters the system of record at scale.
- Calendar-based proficiency targets. Punishes low-volume processes for arithmetic they cannot change.
- Reading quiet as stable. Falling tickets with flat workarounds and falling usage is abandonment, not stabilisation.
What this is for
These three measures exist to answer one decision: can hypercare end?
Adoption tells you people are using it. Proficiency tells you the work is being done correctly. Stabilisation tells you normal operations can absorb what remains without the project team. All three, plus named business owners for the residual issues, is the honest exit criterion — and the alternative, ending hypercare when the budget does, is covered in the go-live recovery playbook.
Anything reported after go-live that does not contribute to that decision, or to the value question behind it, is worth deleting.
For the wider measurement chain from activity through to value captured, see how to measure the success of a change management programme. For the equivalent question before cutover, see what a readiness dashboard should show before go-live.
Some of this is measurable without asking anyone. Transaction logs record behaviour rather than belief, and the most informative moment is usually where the log and the survey disagree.
Once the numbers move, the next question is whether adoption actually caused the benefit; further out, how to know whether a transformation is truly embedded.
