The readiness score is on the slide. It says 72.

Somebody asks whether 72 is good. There is a pause, and then a discussion about whether the survey was fair, whether the response rate was high enough, and whether the sales team always scores low. Ten minutes later the meeting moves on. Nothing was decided.

The problem is not the number. It is that the number was built as a summary rather than as a signal — and a summary has no obligation to trigger anything.

A readiness score earns its place when it tells leadership where to look and when to act. Here is how to construct one that does.

What a score is actually for

A readiness score does three jobs, and it is worth being clear that diagnosis is not one of them.

What it cannot do is explain itself. A score tells you that supply chain is at 61 and falling; it does not tell you why, and it will not tell you what to do. That is the difference between a measurement and a diagnosis, and expecting the first to behave like the second is how readiness scores end up ignored.

Build it as a search tool with an alarm attached, and it becomes genuinely valuable.

Build on constructs, not questions

The most common construction error is averaging whatever questions ended up in the survey. That produces a number with no theory behind it, which means no sub-score points at anything you can actually change.

Start from validated constructs. Holt, Armenakis, Feild and Harris’s readiness for organisational change scale resolves readiness into four beliefs — appropriateness, management support, change-specific efficacy and personal valence. Those four are useful precisely because each maps to a different intervention: a low appropriateness score is a case-for-change problem, low management support is a sponsor and manager problem, low efficacy is a capability and practice problem.

Whatever construct set you use, apply one test to each sub-scale: if this score were low, would we know who acts and what they do? If not, the sub-scale is decoration.

It also helps to keep a small set of items constant across every wave, with stage-specific items layered on top. Strategic questions six months out and operational questions three weeks out are necessarily different, but without a stable core you cannot compare wave to wave — and movement is the score’s most valuable output. This is the logic behind a staged readiness assessment tracked across transformation waves.

Normalise so the number means something

Likert means are poor executive communication. The difference between 3.6 and 3.9 looks trivial on a five-point scale and is, in fact, substantial.

Convert to a 0–100 index:

Score = ((mean − 1) / 4) × 100

A mean of 3.68 becomes 67. Now 65 against 72 reads as the gap it actually is, and the scale runs from “nobody agrees” to “everybody strongly agrees” rather than from an arbitrary 1.

Two disciplines come with this. Reverse-score negatively worded items before aggregating, so higher always means more ready — forgetting this is the most common arithmetic error in readiness reporting. And normalise before you aggregate, not after. The OECD and the European Commission’s Joint Research Centre set out the standard treatment in their handbook on constructing composite indicators, which is worth knowing exists: building an index is a methodological discipline with established rules, not a spreadsheet convenience.

The compensability trap

This is the most dangerous property of any composite score, and the one most readiness reporting walks straight into.

When you average sub-scores, strong components compensate for weak ones. The OECD/JRC handbook treats compensability as a deliberate design decision to be made and declared, not an accident to be discovered later.

In readiness work the consequence is concrete. A function reporting an overall 67 might be carrying strong leadership trust, decent commitment — and training effectiveness at 54. The composite looks survivable. The component is the thing that will hurt you at go-live, because no amount of trust in leadership compensates for people who cannot complete the task.

Two rules follow, and they cost nothing:

Weighting is a judgement. Declare it.

Every composite embeds a weighting, including the ones that look neutral. Averaging everything equally is not the absence of a choice; it is the assertion that all components matter equally, which is rarely what anyone believes.

The handbook’s guidance is that weights should follow the underlying theoretical framework, and that robustness should be tested rather than assumed. Translated into practice: write down your weighting and the reason for it, then check whether a different reasonable weighting would change any decision you are about to take. If it would, the score is not robust enough to carry that decision on its own and you should say so.

A defensible and simple approach is a stable core weighted evenly against stage-specific items, so half the score is comparable across every wave and half reflects where the programme currently is. What matters less is which scheme you pick, and much more that it is explicit and unchanged between waves — changing the weighting mid-programme destroys the trend, which was the point.

Report the spread, not just the mean

Bryan Weiner’s theory of organisational readiness for change argues that readiness is a shared property: shared resolve to implement, and shared belief in collective capability. That has a direct measurement consequence.

A team at 68 might be uniformly moderate, or split down the middle between people at 90 and people at 45. Those are different situations needing different action, and they produce the same score. Under Weiner’s framing the second is arguably not “moderately ready” at all, because the sharedness that constitutes readiness is absent.

So report distribution alongside the mean. The percentage of respondents below a defined threshold is usually the most useful single addition: “Supply Chain 67, but 31% of respondents scored under 50” changes the conversation immediately, at no analytical cost.

Segment properly — and protect small groups

An enterprise-level readiness score is almost useless. Averaging across everyone guarantees that the groups in trouble are cancelled out by the groups that are fine.

Collect the segmentation fields with the survey — function, business unit, country, location, team, management level, role or user group, and wave. Retrofitting them later is impossible without breaking anonymity.

Then set a minimum cell size and hold it. Do not report a subgroup with fewer than about five responses. This protects anonymity, which protects honesty in the next wave, and it also stops the programme reacting to noise: with four respondents, one person having a bad morning moves the score several points.

Publish the rule before the results. Announcing a suppression threshold after someone asks for a breakdown of their eight-person team looks like concealment.

Turning the score into a signal

Everything above produces a good number. A number becomes a signal only when three things are attached to it in advance:

Written as an if-then: “If any function’s operational confidence falls below 60 within six weeks of go-live, the programme director convenes a targeted practice intervention for that function within five working days.”

Agreeing thresholds before the data arrives is the critical detail. Afterwards, every threshold becomes negotiable, and the discussion turns into whether the number is fair rather than what to do about it. It is also the same specificity that makes a 30/60/90-day action plan work rather than drift.

Trend beats level

Nobody knows whether 72 is good. There is no external benchmark that means anything across different organisations, changes and cultures, and comparisons to other companies’ scores are mostly noise.

What is interpretable is your own movement. A 67 that was 74 last wave is a live problem. A 67 that was 58 is a programme working. The same number, opposite meanings — which is why the change since the previous wave belongs on the slide next to every score, and why a one-off readiness survey delivers a fraction of the value of a repeated one.

Minimum reporting standard

Every readiness score report should carry these, without exception:

ElementWhy it is mandatory
Overall scoreThe headline, normalised to 0–100
Valid respondents and response rateLow response rates make the score unrepresentative and are rarely stated
Change since previous waveThe most interpretable figure on the page
Organisational averageThe only benchmark that legitimately applies
Rank among peer groupsLocates who needs attention first
Lowest and highest factorDefeats compensability; points at the intervention
Spread or % below thresholdExposes splits that the mean conceals
Suppressed groupsShows what is not being reported and why

Keep raw responses, scored data and dashboard outputs in separate layers, so any published figure can be traced back to the responses behind it. Readiness scores get challenged — usually by whoever scored lowest — and an auditable trail settles it in minutes rather than undermining the whole instrument.

Readiness measurement Building a readiness score people will act on? Book a 20-minute scoping call to work through constructs, weighting, thresholds and the reporting that turns a number into a decision.
Book a 20-minute scoping call

One warning: do not make it a target

The moment a readiness score appears in someone’s objectives, it stops measuring readiness and starts measuring the effort put into the score.

Charles Goodhart made the original observation in 1975 about statistical regularities collapsing under the pressure of being used for control; the anthropologist Marilyn Strathern later gave it the phrasing everyone quotes — when a measure becomes a target, it ceases to be a good measure.

In readiness measurement this shows up quickly: managers coaching teams on how to answer, surveys quietly timed for good weeks, uncomfortable groups omitted from distribution. All of it is rational once the score is a performance measure, and all of it destroys the only thing the score was for.

Keep the score as an instrument, not an objective. Hold managers accountable for acting on what it shows — which is where manager reinforcement genuinely matters — never for the number itself.

The test of a good score

It is not whether the methodology is elegant. It is whether, when the number moves, somebody does something.

If your last readiness report changed nobody’s behaviour, the fix is usually not a better survey. It is thresholds agreed in advance, owners named against them, the lowest factor published alongside the headline, and the spread shown next to the mean.

None of that requires more data. It requires deciding, before the results arrive, what would make you act — which is the difference between a readiness score and a readiness signal.

For turning the resulting signal into owned actions, see from readiness survey to leadership action and designing metrics that create decisions rather than decorative dashboards. The change readiness resources cover the wider measurement picture.

Ritvars Mētra

Ritvars Mētra

Founder of ReadinessCompass

Ritvars Mētra is the founder of ReadinessCompass, where he develops practical tools for understanding and managing organisational change complexity. His work focuses on adoption readiness, stakeholder analysis, and evidence-based change management for large-scale software and AI implementations.

View full profile →