The readiness pack says Manager readiness: amber.
Underneath it is a survey that managers completed about themselves, asking whether they feel prepared to support their teams through the change. Seventy-one per cent said yes.
That number is close to meaningless, and the reason is well documented. It is also fixable without much effort, because the things worth measuring here are mostly observable.
Reinforcement is a behaviour, not a feeling
Before measuring anything, define what you are measuring. “Supporting the change” is not measurable. The specific recurring acts that constitute reinforcement are.
In Prosci’s ADKAR model, reinforcement is the fifth element — what sustains a change after people are able to perform it. At manager level that resolves into a short list of observable behaviours:
- Asking for the new artefact rather than the old one.
- Reviewing exceptions or errors from the new process with the team.
- Protecting time for people to practise.
- Challenging reversion when someone uses the old route.
- Escalating genuine blockers rather than absorbing them locally.
- Acknowledging the difficulty honestly instead of defending a broken step.
Every one of those is something a person either does or does not do in a given week, and that is what makes them measurable. If you cannot name the behaviours, you cannot measure readiness for them — you can only measure sentiment about them.
Why asking managers about themselves fails
There is a precise, long-standing finding here.
Harris and Schaubroeck’s meta-analysis of self, peer and supervisor ratings in Personnel Psychology found that self-ratings correlate only moderately with supervisor ratings (ρ = .35) and with peer ratings (ρ = .36) — while peer and supervisor ratings correlate substantially with each other (ρ = .62).
Read that carefully, because it is stronger than “self-report is unreliable.” Two independent external sources agree with each other far more than either agrees with the person being rated. The outlier is the self-rating.
Add the incentives specific to this situation. A manager asked whether they are ready to lead their team through a change is being asked, in front of their own management line, to assess their own competence at a core part of their job. Very few will report that they are not. The instrument is measuring something, but it is not reinforcement.
The conclusion is not to stop asking managers anything. It is to stop scoring them on their own answers.
Ask the team instead
A neighbouring field solved this measurement problem decades ago.
Safety research faced exactly the question of how to tell whether a supervisor genuinely reinforces something day to day, as opposed to endorsing it when asked. Dov Zohar’s work on leadership, safety climate and outcomes in work groups treats group climate as emerging from patterns of supervisory practice rather than isolated statements, and the field measures it by asking the group what their supervisor actually does — aggregated to team level.
The transfer is direct. Add four to six items about manager behaviour to the survey you are already sending to end users, and aggregate them to team level. You get a reinforcement measure from the only people with daily evidence, at no extra survey cost, and you avoid the self-rating problem entirely.
Two design cautions. Keep the items behavioural and recent — “in the last two weeks, my manager reviewed…”, not “my manager supports the change.” And respect a reporting floor: a manager with four reports cannot be reported individually without destroying anonymity, so aggregate to a level or a cluster and say so before you collect anything.
Four layers, and which ones to score
Manager reinforcement readiness sits in four distinguishable layers. Most programmes measure the second and report it as though it were the third.
| Layer | Question | Source | Use it to |
|---|---|---|---|
| 1. Precondition | Do they have what they need? | Administrative records — briefing dates, agreed hours, training completion | Explain a low score; fix the cause |
| 2. Intention | Do they say they will? | Manager self-report | Diagnose only — never score |
| 3. Behaviour | Are they doing it? | Team-reported items, plus observable artefacts | Score this |
| 4. Effect | Is it working in their team? | Adoption, unaided completion, workaround closure by team | Score this |
The rule that follows is simple: score layers 3 and 4; use layers 1 and 2 to explain them. A team-level behaviour score of 42 tells you there is a problem. The precondition layer tells you whether it is because the manager was briefed two days before go-live and given no hours — which is a programme failure, not a manager failure.
Items that work
Team-reported (in the end-user survey, five-point agreement, aggregated to team):
- In the last two weeks, my manager has asked about how the new process is going for me.
- My manager uses the new reports and information in our team meetings.
- My manager has made time available for me to practise or ask questions.
- When I raise a problem with the new way of working, something happens.
- My manager is clear that the old way is no longer acceptable.
The fourth is the most diagnostic item in the set. It measures whether raising a problem has ever produced a result, which is the single best predictor of whether people keep reporting problems at all.
Manager-reported (diagnostic only, never scored):
- How many hours a week do you currently have for this? (a number, not a scale)
- What question from your team can you not answer?
- What would you need to make this easier?
Open questions rather than ratings. You are not scoring the manager; you are collecting the cause of whatever the team-reported score turns out to be.
The indicators that need no survey
The strongest measures here are administrative and free, and they can be collected before any survey runs.
- Was the manager briefed before their team? Check the calendar, not the intent.
- Are the agreed hours real? Known to whoever owns their operational targets, or not agreed at all.
- Have they used the escalation route? A manager who has never escalated in eight weeks of a difficult go-live is either exceptionally lucky or absorbing problems silently.
- Is the old report still being requested? The cleanest reinforcement signal available, and it comes from the reporting system rather than from anyone’s opinion.
- Do they attend the manager forum? Attendance is a weak proxy for commitment but a strong proxy for whether the programme has any route into that team.
- Can they complete a core task unaided? Observed, not asked — the same distinction that makes UAT a capability instrument rather than a confidence one.
“Is the old report still being requested” deserves particular attention. It is objective, continuously available, and it measures the behaviour that matters most — because as long as the old artefact is asked for, it will be produced, and the new way remains optional.
Reading the result
Three rules for interpreting what comes back.
Compare across teams, not against an absolute. Nobody knows whether a reinforcement score of 64 is good. What is interpretable is that one function sits at 64 while comparable teams sit at 78 — and that the gap is stable across two waves.
The self-versus-team gap is itself the finding. Where you have both, plot them together. A manager reporting high confidence whose team reports low reinforcement is the highest-risk cell on the grid, because nobody in that chain currently believes there is a problem. That is the same confidence-versus-capability gap that shows up wherever self-report is used without observed evidence beside it.
Treat a persistent middling score as worse than a low one. A team sitting at 58 for three consecutive waves has stopped being a finding and become background, which is exactly how known problems stop being seen.
When the score is low
The reflex is a manager engagement session. Check the precondition layer first, because the cause is usually structural and the session will not touch it.
Baldwin and Ford’s review of transfer of training identifies the work environment — supervisory support and the opportunity to apply what was learned — as one of three determinants of whether anything transfers to the job. Managers are that work environment for their teams. But they are also inside one themselves, and if their own environment gives them no time, no answers and no authority, reinforcement is not something willingness can supply.
So map a low score against the five preconditions in what line managers need before go-live: information ahead of their team, answers to personal-impact questions, capability in the change, time, and authority. In most cases one of those is missing, and it was missing by design.
One further check: reinforcement requires something to reinforce. If a manager is telling their team to follow a process that genuinely does not fit a recurring case, and nobody with authority has ruled on it, the failure sits with the missing decision route, not the manager.
Cadence
Administrative indicators can run continuously and cost nothing. Team-reported items belong in whatever readiness or adoption survey is already going out — once before go-live to set a baseline, then at the same cadence as the rest of the instrument during hypercare.
Do not run a separate manager survey unless you have a specific question that the team-reported items cannot answer. A dedicated survey adds burden, invites the self-rating problem back in, and signals to managers that they are being assessed rather than supported — which reliably degrades the honesty of everything else they tell you.
The short version
Name the behaviours. Ask the team, not the manager. Collect the free administrative indicators before running any survey at all. Score behaviour and effect; use preconditions and manager self-report to explain what you find. Compare teams against each other rather than an invented benchmark. And treat the gap between what a manager believes and what their team reports as a finding in its own right.
None of that requires a new instrument. It mostly requires not scoring people on their own account of themselves — which is the one thing almost every manager readiness measure currently does.
Measuring reinforcement is only useful if something can be done about a poor result. Most of what actually reinforces is not encouragement but the removal of obstacles — making the new way the easy way rather than adding pressure to choose the harder one.
For why this layer determines adoption, see the manager as adoption multiplier. For putting the resulting measure somewhere useful, see what a readiness dashboard should show before go-live.
