The slide says two things. Adoption reached 82%. Order-to-cash cycle time fell 19%.
Both numbers are true. Between them sits an arrow that nobody has examined, and the arrow is the entire claim.
When a CFO asks whether the improvement was actually caused by the change programme, the honest answer in most organisations is that nobody checked. Adoption went up, a number moved, and the two were placed next to each other. That is a correlation presented as a result — and it is the weakest link in almost every benefits claim made after go-live.
It is also more fixable than it looks, mostly with decisions made before cutover.
Five ways a before-and-after comparison lies to you
Evaluation methodology has catalogued these for sixty years. Shadish, Cook and Campbell’s work on quasi-experimental design sets out the standard threats to internal validity — the named ways an apparent cause-and-effect relationship turns out to be something else. Five of them show up constantly in transformation benefit claims.
| Threat | What it looks like in a programme |
|---|---|
| History | Something else changed in the same window — a competitor exited, a pricing change landed, a backlog was cleared, a reorganisation removed a handoff |
| Maturation | The team would have improved anyway. Staff who were new last year are experienced now, and that trend predates your project |
| Regression to the mean | You intervened because performance was unusually bad. Extremes drift back toward average on their own, with or without a programme |
| Selection | The pilot site volunteered, so it was already better led, better staffed or more motivated than the sites you compared it to |
| Instrumentation | The new system measures the thing differently from the old one, so part of the “improvement” is a definition change |
Two of those deserve more attention than they get.
Regression to the mean is the trap almost nobody names, and transformation programmes are structurally exposed to it. Programmes are commissioned when performance is bad — that is what triggers the investment. Bad periods are frequently unusual periods, and unusual periods are followed by more ordinary ones regardless of what anyone does. A meaningful share of post-go-live improvement in poorly performing units is the world returning to normal.
Instrumentation is the one specific to system replacements, and it is embarrassing when it surfaces late. If the legacy system started the clock at order confirmation and the new one starts it at order entry, cycle time will improve on go-live day without anybody doing anything differently. Before claiming a movement, confirm the metric definition survived the migration unchanged. Surprisingly often it did not.
You probably already have a comparison group
Here is the most useful thing in this article, and almost no programme uses it.
If you are rolling out in waves, or across multiple sites or countries, the units that have not gone live yet are a control group. They are experiencing the same market, the same seasonality, the same economic conditions and the same organisational noise — but not your change. That is a natural experiment, available for free, and it disappears the moment the last wave goes live.
The comparison is simple. Do not compare wave 1 after against wave 1 before. Compare the change in wave 1 against the change in not-yet-live wave 2 over the same period.
If cycle time fell 19% in live sites and 4% in not-yet-live sites, your programme plausibly accounts for something like the difference — and the 4% was going to happen anyway. That single subtraction removes history, maturation and most seasonality in one step, because both groups were exposed to them equally.
Three conditions make it credible: a baseline measured on both groups before wave 1 goes live; broadly parallel trends between the groups beforehand; and an identical metric definition on both sides. All three are cheap if arranged in advance and impossible to retrofit afterwards.
When there is no comparison group
Single-site big-bang deployments have no untreated units. Two options remain, and both are worth having.
Use the trend, not the point. With enough pre-go-live data points — twelve monthly observations rather than one “baseline” — you can establish both the existing trajectory and its normal variability. A post-go-live movement only counts as evidence if it breaks the pre-existing trend by more than the ordinary variation. One before-figure and one after-figure cannot distinguish a real shift from a noisy month, which is the same common-cause problem that makes people over-react to small dashboard movements.
Look for a dose-response gradient. This is the practical workhorse and it needs no control group at all. If adoption caused the benefit, then units with more adoption should show more benefit. Plot adoption against outcome across teams, sites or functions and look at the slope.
A clear gradient is real evidence: the mechanism behaves the way the theory says it should. A flat line is more informative still — if the high-adoption and low-adoption units improved equally, then whatever produced the improvement, it was not adoption. That finding is unwelcome and much better learned internally than in front of a CFO.
One prerequisite: adoption and outcome must be measured at the same unit of analysis. Adoption by individual and cycle time by site cannot be related to each other. Deciding that both will be captured by site, or both by team, is a five-minute decision before go-live and an unfixable problem afterwards.
Contribution, not attribution
For most transformation programmes a clean causal estimate is simply not available. Several things changed at once, on purpose, and nobody was ever going to randomise sites.
Evaluation has a recognised answer for exactly this situation. John Mayne’s contribution analysis, in the journal Evaluation, was developed for interventions where attribution cannot be established but a credible causal claim is still needed. Rather than asserting that the programme caused the outcome, it builds an evidenced argument that the programme contributed, and explicitly confronts the alternatives.
In practitioner form:
- State the causal chain explicitly. Adoption of the new credit-check step reduces disputed invoices, which reduces rework, which shortens cycle time. Write it down as links, not as a slogan.
- Identify the assumption in each link — the thing that must be true for it to hold.
- Gather evidence on the intermediate links, not only the endpoints. If disputed invoices did not fall, the chain is broken in the middle whatever the cycle time did.
- List the alternative explanations and test each one, rather than waiting for someone else to raise them.
- Say what you cannot rule out. The claim gets stronger, not weaker, for naming its own limits.
The middle-link evidence is the part most programmes skip, and it is the cheapest thing on the list. Intermediate measures are usually already collected; nobody thought to connect them into a chain.
What a defensible claim sounds like
| Weak | Defensible |
|---|---|
| Adoption reached 82% and cycle time fell 19%, so the programme delivered | Cycle time fell 19% in live sites against 4% in not-yet-live sites over the same period |
| Benefits are on track | Within live sites, the three highest-adoption sites improved 26% against 9% for the three lowest |
| The new process is working well | Disputed invoices — the intermediate step in the causal chain — fell 31%, consistent with the mechanism |
| Improvement is due to the transformation | The metric definition was held constant across the migration and independently checked |
| — | The main alternative explanation, the Q3 volume decline, affected both groups equally and does not account for the gap |
The right-hand column is not more work. It is the same data, arranged so that it answers the question actually being asked.
Why the honest version is worth more
There is a temptation to make the largest claim the data will tolerate, and a structural reason it is a bad trade.
Bent Flyvbjerg’s work on major projects separates optimism bias — the cognitive tendency to judge outcomes more favourably than experience warrants — from strategic misrepresentation, the deliberate overstatement of benefits to secure and retain approval, with reference class forecasting proposed as the corrective. Finance functions have usually met both before. A benefits claim with no stated limitations reads as one of them.
A smaller claim that survives challenge does more for the next business case than a larger one that collapses under a single question. And an organisation that can genuinely tell which of its changes produced value can stop funding the ones that do not — which is worth considerably more than any individual programme’s reported number.
Four decisions to take before go-live
All of this depends on choices made while attribution is still possible. After cutover, most of it cannot be recovered.
- Baseline both groups, not just the first wave. Twelve pre-period observations where you can get them, on live and not-yet-live units alike.
- Freeze and document the metric definition, and check after migration that it still means what it meant. Record the definition somewhere durable, not in a deck.
- Fix the unit of analysis so adoption and outcome are captured at the same level.
- Write down the causal chain and its intermediate measures while the business case is being argued, when everyone is still clear about why the benefit was expected.
Those four are a morning’s work during planning. Attempted a year later, they are a research project with missing data.
The question behind the question
When someone asks whether adoption caused the benefit, they are rarely trying to catch anyone out. They are asking whether the organisation should do more of this.
That question deserves a real answer, and a real answer has a shape: here is the movement, here is the comparison group, here is the gradient across adoption levels, here is the intermediate link behaving as predicted, here is what else was happening, and here is what we cannot rule out.
Anything shorter is two numbers and an arrow.
For the measurement chain this sits at the end of, see how to measure the success of a change management programme. For the adoption and proficiency measures that feed it, what to measure after go-live; and for what quietly consumes the benefit in between, why benefits leak after go-live.
