Most change readiness surveys are written the same way. Someone opens a document, lists the things that feel important, turns each into a statement, adds a five-point agreement scale, and sends it out. Forty-two questions, three weeks before go-live.
The result is usually a set of numbers in the high threes that nobody can act on, and a quiet decision never to run one again. That outcome is not bad luck. It is what this design produces.
Start from the construct, not the topic list
The first question is what you are trying to measure, and “readiness” is not yet an answer.
Bryan Weiner’s theory of organizational readiness for change splits it into two facets: change commitment — whether people are resolved to implement the change — and change efficacy, the shared belief in their collective capability to do it, shaped by task demands, resource availability and situational factors.
That distinction does real work in survey design, because the two facets have completely different remedies. Low commitment is a case-for-change and involvement problem. Low efficacy is a capacity, resourcing and capability problem. A survey that returns a single readiness number cannot tell you which you have, which is why so many of them lead to a communications plan regardless of the underlying issue.
Do not invent items when validated ones exist
Most organisations write readiness items from scratch. There is a tested alternative: the Organizational Readiness for Implementing Change measure, developed from Weiner’s theory and psychometrically assessed by Christopher Shea and colleagues in Implementation Science. It is short, open access, and built specifically around commitment and efficacy as separate factors.
Using validated items has two advantages beyond saving an afternoon. The wording has been tested for how people actually interpret it, and the factor structure means the responses can be analysed as two constructs rather than averaged into mush.
It was developed in healthcare implementation research, so some adaptation is needed. Adapt the context and keep the structure; the temptation to rewrite every item into house language is how a validated instrument becomes an invented one.
Length is not a courtesy question
The forty-two-question survey does not merely annoy people. It systematically corrupts its own data.
Jon Krosnick’s work on response strategies for coping with the cognitive demands of attitude measures, in Applied Cognitive Psychology, describes what happens when answering properly requires more effort than a respondent is willing to spend: rather than stopping, they satisfice — producing a satisfactory answer instead of an optimal one. The strategies he identifies are all visible in readiness data:
| Satisficing strategy | How it appears in your results |
|---|---|
| Choosing the first reasonable option | Clustering on the first plausible scale point rather than the accurate one |
| Agreeing with the assertion in the question | Uniformly positive scores on positively worded items — acquiescence, not readiness |
| Endorsing the status quo | Systematic under-reporting of appetite for change |
| Failing to differentiate across items | Straight lines down the page. The single most common defect in long readiness surveys |
| Saying “don’t know” | Missing data concentrated in exactly the groups you most need to hear from |
| Choosing at random | Noise that looks like moderate readiness |
Note that satisficing does not produce obviously bad data. It produces plausible, mid-range, low-variance data — which is exactly what a readiness survey that has failed looks like. If every function scores between 3.4 and 3.8, the most likely explanation is not that readiness is uniform.
Write the decision before the questions
The discipline that removes most of the length problem is deciding, for each candidate item, what you would do differently depending on the answer.
Take a real example. “I understand why the organisation is making this change.” If it comes back at 40 per cent, what happens? If the honest answer is “more communication”, and more communication is already planned regardless, the item is not informing a decision. It is producing a number.
Applied honestly, this test usually removes half the survey. What survives tends to be items about capability, capacity and specific blockers — because those are the ones where a bad answer forces a genuine choice about scope, date or resourcing. That is the same logic that separates a decision dashboard from a decorative one, applied one step earlier.
Ask about behaviour where you can
Attitude items are cheap to write and weak to act on. Where a behavioural or factual item is available, it is worth several agreement scales:
- Instead of “I feel prepared for the new process”, ask “Have you completed a full order in the new system without help? Yes / No / Have not tried”.
- Instead of “My manager supports the change”, ask “In the last two weeks, has your manager discussed the change with you? Yes / No”.
- Instead of “I have capacity to absorb this change”, ask “What will you stop doing to make time for it?” as free text.
These are harder to satisfice, harder to misinterpret, and produce findings a sponsor can act on the same week. The third one in particular is worth more than any scale item: a page of “nothing” answers is the clearest capacity finding you will ever obtain.
A workable shape
- Ten to fifteen items. Not forty. Length buys satisficing, not insight.
- Commitment and efficacy reported separately, never averaged into one score.
- Two or three behavioural items that can be verified independently.
- One free-text question that could return something you did not expect. “What is the one thing that would make this easier?” earns its place.
- Segmentation captured at the start — function, location, role — because an organisational average conceals everything that matters.
- A stated decision rule, written before the data arrives, for what result would change the plan.
The last item is the one most often skipped and the one that determines whether the exercise was worth running. A survey with no pre-agreed threshold produces a discussion; a survey with one produces a decision.
One design decision sits outside the instrument itself: how often to run it. Cadence should follow the decisions the data serves rather than the reporting calendar, and measuring faster than a construct actually moves produces noise the programme will then over-interpret.
More on turning readiness data into leadership action in the Change Readiness Hub, or see how a Readiness Diagnostic Sprint gathers this evidence when there is no time for a survey cycle.
