Here is a business readiness criterion, taken almost verbatim from a real programme: 95% of impacted users have completed training.
It looks rigorous. It has a number, a population and a verb. It will be reported as met, because it is designed to be met — and it permits a state of the world in which 95 per cent of users have clicked through an e-learning module, retained almost none of it, and cannot complete a single end-to-end task without help.
The criterion is not wrong about anything. It simply measures whether people were processed, and then the programme treats that as evidence that people are ready. Those are different claims, and the gap between them is where most go-live disappointment lives.
Two well-documented reasons criteria decay
This is not a failure of diligence, and writing more careful criteria is only part of the answer. There are two specific mechanisms at work, both studied outside change management, and knowing them changes how you write.
The measure attracts pressure. In 1979 the psychologist Donald T. Campbell, writing on the evaluation of social programmes, stated what is now known as Campbell’s Law: the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures, and the more apt it will be to distort and corrupt the processes it is intended to monitor.
A training completion figure that nobody looks at is a reasonable proxy for effort. The moment it becomes a gate on a go-live date, it stops being a proxy for anything, because the fastest route to 95 per cent is to make the module shorter and the questions easier. Nobody involved is acting in bad faith. The indicator is simply now load-bearing, and load-bearing indicators bend.
The measure replaces the thing it stood for. The second mechanism is subtler and more damaging. In The Accounting Review, Willie Choi, Gary Hecht and William Tayler named it surrogation: managers fail to fully appreciate that measures are only representations of the underlying construct, and begin to act as though the measure is the construct.
Applied here, the construct is “the business can operate on Monday”. The measure is training completion. Surrogation is what has happened when a programme stops asking whether the business can operate and starts asking whether training completion has hit 95 — and it is why a room full of intelligent people can watch a criterion be met and a go-live fail, without anyone noticing a contradiction.
Their research found surrogation is worst when people are held to a single measure of a construct, and less pronounced when several measures of the same construct are used. That finding is directly actionable, and it is the single most useful design rule in this article.
The anatomy of a criterion that holds
A readiness criterion needs six parts. Most have two.
- The construct — what you actually care about, in plain words. “Order entry clerks can process a standard order without help.”
- The measure — the observable stand-in. Unaided completion of three defined scenarios.
- The evidence — what physically gets produced, and by whom. An observation sheet per person, completed by someone who does not report to the process owner.
- The threshold — agreed before the data exists.
- The owner — the person who asserts it and will still be there in six weeks.
- Compensability — whether strength elsewhere can offset a miss here, stated explicitly.
The separation of measure from evidence is the part that does most of the work. “Training completion 95%” conflates them: the measure is the evidence, produced by the system that has an interest in the number. Once you must name who produces the evidence and how, weak criteria become visibly weak on the page.
| Fudgeable | Harder to fudge |
|---|---|
| 95% of users trained | Every role has at least three people who completed the core scenario unaided, timed, in the last three weeks |
| Cutover plan approved | Cutover rehearsed end to end; the two failures found have owners and retest dates |
| Data migration signed off | The data owner in each function has reviewed a sample of their own records and confirmed the content is usable |
| Support model in place | Named individuals released in writing by their managers, with rota published and cover agreed for the first two weekends |
| Business ready for go-live | Each receiving manager has stated in writing what their team will stop doing in weeks one to three |
Notice what the right-hand column has in common. Each one requires a specific named human to say something specific that they can be held to, and none can be produced by a report. That is the difference, and it is uncomfortable by design.
Steal the two-column structure from government assurance
There is a ready-made format for this, and it has been public for twenty years. The Gateway review workbooks — for example Gate 4: Readiness for service — are built as two columns: areas to probe on the left, evidence expected on the right.
So the question “Is the organisation ready for business change?” is not answered with a percentage. The evidence expected is listed: agreed plans for business preparation and transition, a documented communications plan, staff trained and informed, a clearly defined service management function in place. And the companion question — can the organisation implement the new services and maintain existing services — expects a resource plan showing capacity and capability, not a statement of intent.
Writing criteria in two columns forces the separation that matters. A criterion whose evidence column reads “confirmation from the programme” is not a criterion. It is a request for reassurance.
Six tests to run on every criterion you write
- The cheat test. What is the cheapest way to satisfy this without the underlying thing being true? If a plausible answer exists, that is what will happen — not through dishonesty, but because programmes under time pressure take the cheapest available route to green.
- The independence test. Who produces the evidence, and do they benefit from the answer? Evidence produced solely by the party being assessed is not evidence.
- The falsifiability test. Could this criterion fail? If there is no realistic state of the world in which it is not met, it is decoration.
- The recency test. When was the evidence generated? A capability test from three months and two configuration changes ago is a historical document.
- The population test. Whose capability was measured — volunteers and super users, or the median performer and the person who was on leave during training? The second group defines day one.
- The consequence test. If this is missed by a small margin, what actually happens? A criterion with no consequence attached will be waived, and everyone knows it before the meeting starts.
The cheat test is the one that changes the most drafts. Run it honestly on “95% trained” and the answer arrives in about four seconds.
Three structural defences
Use more than one measure per construct. This follows directly from the surrogation research. If “people can do the work” is carried by training completion alone, that number becomes the goal. Carry it with three — unaided completion rate, time per transaction against baseline, and the volume of expert interventions during rehearsal — and no single number can stand in for the construct. Gaming three measures that disagree with each other is considerably harder than gaming one. The same logic is what makes a composite readiness score more robust than any of its inputs.
Involve the receiving managers in choosing the criteria. In a follow-up study, the same researchers found that involving managers in the selection of a strategy reduced their tendency to surrogate — while merely involving them in deliberation did not. That distinction is worth taking literally. Asking process owners for comments on criteria the programme has already written will not change behaviour. Asking them to choose, from a set of options, how their own function’s readiness should be judged does. It also produces better criteria, because they know which corner their team would cut.
Set thresholds before the data exists. A threshold agreed two weeks before go-live is negotiated against a known answer, and it always lands just below wherever the number happens to be. Agreed three months out, it is a genuine test. Write the date the threshold was set next to the threshold; it is remarkable how much that one habit disciplines the conversation.
Say which criteria cannot be traded
Most readiness reporting is implicitly compensatory: strong scores in six areas visually outweigh a weak one in the seventh, and the eye reads the overall picture as acceptable.
Operationally this is false. A first-class finance function does not compensate for a warehouse that cannot pick, because the order still does not ship. Some criteria are floors, not contributors, and the time to say so is when they are written — not during the meeting where one of them is amber and everyone is looking at the average.
Mark each criterion as compensable or not. Typically three or four are floors: capability in the highest-volume customer-facing process, the ability to invoice, statutory or regulatory obligations, and the support model. Everything else can be carried with mitigation.
When something genuinely cannot be measured
Some things that matter resist measurement. Whether managers will reinforce the change. Whether people trust the data. Whether the organisation has the appetite for one more system after the last three.
The wrong response is to invent a number, because an invented number attracts exactly the corruption pressure Campbell described and then substitutes for the thing it was standing in for. The right response is a stated judgement with a named owner and its basis recorded: “The operations director judges her managers are not yet able to answer their teams’ questions, based on the two briefings held in August. Mitigation: a manager-readiness session in week 40, retested in week 42.”
That is auditable, challengeable and honest about its own basis. It is far better than a fabricated 78 per cent, and it is how the parts of readiness that live in manager capability should be handled.
The test of a good criterion
A good business readiness criterion has one property above all others: satisfying it requires doing the thing you actually wanted. If there is a cheaper way to turn it green, that way will be found, and the programme will arrive at go-live with a full set of met criteria and an organisation that cannot work.
Criteria written this way are unpopular while they are being drafted, because they commit named people to specific claims and they can fail. That is precisely what makes them worth having — and it is what turns a go/no-go meeting into a decision rather than a ratification.
More on turning readiness evidence into leadership decisions in the Change Readiness Hub, or see how a Readiness Diagnostic Sprint builds criteria and gathers the evidence behind them in two to three weeks.
