Escalations rarely arrive as surprises. They arrive as surprises to the people who could have acted.

By the time a readiness problem reaches a steering committee as a formal escalation, it has usually been present for two measurement waves, visible to the affected team for longer than that, and mentioned at least once in a meeting where it did not survive the summary. The information existed. The detection failed.

A readiness heat map is a detection instrument. Its value is not that it displays data attractively — it is that it shortens the gap between a problem existing and a problem being noticed. This is about how to read one for that purpose: which shapes matter, which movements are noise, and which cell on the grid is the most dangerous.

(If you need the construction rules first — encoding, colour, small multiples — those are covered separately in how to visualise readiness by function, country and wave.)

Escalation is a detection failure

An escalation is simply a risk that became loud enough to force attention. Loudness is not correlated with importance — it is correlated with how much damage has already been done, and with whether someone was willing to make a fuss.

That second condition is the problem. Information systems research calls it the mum effect: the reluctance to pass on unwelcome messages, recognised as a contributor to project failure. Bad news gets softened at every reporting layer until what arrives at the top no longer sounds like a reason to act.

A heat map routes around that, and this is the whole argument for it. It is mechanical. Nobody has to be brave for a cell to go dark. The grid reports the same number whether or not anyone wants to discuss it, which converts “raising a concern” from an act of individual courage into a routine reading of a chart.

Four shapes worth recognising

Most people read a heat map cell by cell, looking for the darkest square. That finds the worst current score, which is rarely the most important finding. The information is in the shape.

Illustrative patterns Four shapes and what each one means
1. The falling row
W1W2W3
767877
726351
747375
One function losing ground while the others hold. A local cause, and it is compounding.
2. The dark column
W1W2W3
767259
747056
777361
Everyone drops at the same wave. The cause is programme-level — do not fix it locally.
3. The isolated cell
W1W2W3
767475
737148
757674
A sudden single-cell drop. Something specific happened, recently, to one group.
4. The stuck cell
W1W2W3
767778
575857
747576
Never moves, never escalates, stops being discussed. The most dangerous cell on the grid.

Rows are functions, columns are measurement waves. Darker means lower readiness; every cell carries its number so colour locates the pattern rather than carrying the value.

The falling row

A single function declining across waves while its peers hold steady. The cause is local — a manager, a role redesign, a site-specific process, a team carrying more of the change than anyone accounted for.

Direction matters more than level here. A function at 51 and falling is a more urgent conversation than one that has sat at 48 since the start, because the trajectory tells you the situation is still deteriorating.

The dark column

Everyone drops at the same wave. This is the pattern most often misdiagnosed, because it presents as several separate local problems and generates several separate local remediation plans.

It is one problem. Something programme-level happened between waves: training landed badly across the board, a go-live date moved, a system access issue hit everyone, a leadership message did not survive contact. Fixing it function by function is expensive and will not work.

A useful diagnostic reflex: before commissioning three interventions, check whether the dip is column-shaped. If it is, you have one cause.

The isolated cell

One group, one wave, a sharp drop from where it was. This is the highest signal-to-noise pattern on the grid and the one most worth investigating immediately, because a sudden change has a recent, findable cause.

The investigation is usually short. Something happened to that group between the two measurements, and the people in it know exactly what it was. A 5 Why analysis with two or three of them will normally produce the answer inside an hour.

The stuck cell

And then the dangerous one: the cell that has been at 57 for three consecutive waves.

It never moves, so it never appears in a “biggest changes” slide. It is not the worst score, so it never tops a ranking. It has been discussed before and nothing catastrophic happened, so it has stopped feeling like news. Everyone has learned to read past it.

This has a name, and it comes from a serious place.

Why persistent amber stops being seen

Diane Vaughan’s study of the Challenger launch decision produced the concept of the normalisation of deviance: the gradual process by which unacceptable practice becomes acceptable. As the deviation repeats without catastrophic result, it becomes the organisational norm. Each flight with O-ring anomalies made the next anomaly more acceptable. The baseline shifted.

Readiness reporting does this efficiently. A function sits below threshold at wave one; it is flagged, discussed, an action is raised. At wave two it is still below threshold, and someone observes that it was below threshold last time too. By wave three it is context rather than a finding — “supply chain is always around 57” — and the grid has quietly taught everyone that 57 is normal for supply chain.

Then go-live arrives and supply chain cannot operate, and the escalation lands as though nobody saw it coming. It was on the grid the whole time, in the same colour, which is precisely why it became invisible.

The countermeasure is to age the cells. Add one column to the reporting: waves at or below threshold. A function at 57 for three consecutive waves shows a 3, and that number rises visibly while the score does not. It converts persistence — which the eye ignores — into a rising figure, which the eye does not.

It costs one column and it is the single highest-value addition you can make to a readiness heat map.

Noise, signal, and the temptation to tamper

The opposite failure is just as expensive: reacting to movement that means nothing.

Walter Shewhart’s distinction between common-cause and special-cause variation, later central to W. Edwards Deming’s work on knowledge of variation, is the right frame. Common-cause variation is the ordinary noise of the process — who happened to respond, what week it was, how the question landed. Special-cause variation comes from an identifiable outside source. Treating the first as though it were the second is what Deming called tampering, and his funnel demonstration showed that each round of well-intentioned adjustment made the outcome worse.

Readiness programmes tamper constantly. A function moves from 71 to 68 and a remediation plan is commissioned. It moves back to 70 next wave, the plan is credited with the recovery, and the organisation learns a false lesson about what works.

Two practical rules follow:

Between them, these two sections define the actual job: react to the stuck cell that everybody has stopped seeing, and stop reacting to the three-point wobble that everybody notices.

Turning a pattern into a response

Detection only prevents escalation if something follows it. Attach the response to the pattern in advance, before anyone knows the results.

PatternMost likely causePre-agreed response
Falling rowLocal: manager, role change, capacityFunction owner investigates within one week; report cause, not score
Dark columnProgramme-level event between wavesSingle programme-level root cause analysis; no local plans until it is found
Isolated sudden dropA specific recent event in one groupDirect conversation with that group inside 48 hours
Stuck below thresholdAccepted as normal; nobody owns itAutomatic escalation at the third consecutive wave, regardless of level
Movement inside the noise bandCommon-cause variationNo action. Record and move on
Uniformly mid grid, no spreadMeasurement problem, not readinessReview instrument and response rate before interpreting anything

The fourth row is the one that does the work. An automatic trigger on persistence — not severity — is what defeats normalisation of deviance, because it removes the requirement for someone to notice and choose to raise it again.

Agree these before the data arrives. Afterwards, every threshold becomes negotiable, and the meeting turns into a discussion about whether the number is fair. The same logic that makes a 30/60/90-day plan hold applies here: a trigger without a named owner and a defined first move is a preference, not a control.

Adoption risk detection Finding out about adoption risk too late? Book a 20-minute scoping call to set up readiness reporting that surfaces risk while it is still quiet — thresholds, triggers and owners.
Book a 20-minute scoping call

Cadence sets your blind spot

A heat map refreshed quarterly gives you a quarterly blind spot. If a function collapses two weeks after a measurement wave, you will find out ten weeks later — by which time it has escalated through the route you were trying to avoid.

Match the refresh rate to the decision rhythm and to proximity to go-live. Six months out, quarterly is fine. Six weeks out, quarterly is negligent — and this is the window where a shorter, narrower pulse on the two or three functions you are already worried about beats a full enterprise survey nobody has time to run.

Detection latency, not sample size, is what determines whether the instrument prevents escalations or merely documents them afterwards.

What a heat map cannot do

Three limits, worth stating plainly so the instrument is not asked to do work it cannot.

The question to ask of the grid

Not “what is the lowest score?” That one is already being managed, because it is visible and embarrassing.

Ask instead: which cell has been amber longest without anybody doing anything about it?

That is where your next escalation is currently sitting, in plain sight, in a colour everyone has grown used to.

For building the underlying number, see turning survey data into a readiness score leaders act on. The adoption risk resources cover the wider set of signals that predict trouble before go-live.

Ritvars Mētra

Ritvars Mētra

Founder of ReadinessCompass

Ritvars Mētra is the founder of ReadinessCompass, where he develops practical tools for understanding and managing organisational change complexity. His work focuses on adoption readiness, stakeholder analysis, and evidence-based change management for large-scale software and AI implementations.

View full profile →