The launch went ahead. You said it was too early, and you were overruled.

Now the data is wrong in ways nobody predicted, three core workflows do not work end to end, the service desk is drowning, and people who were merely sceptical two weeks ago are now openly angry. Support is running daily office hours and publishing an FAQ that grows by the hour. Someone senior has started using the word “rollback” in meetings. And you are the change person, which means everyone is looking at you to say what happens next.

This is a recognisable situation, and it is more common than the case studies admit. What follows is a sequence for working through it — what to do first, what to deliberately not do, and how to know when the crisis is actually over.

One thing first, because it affects how clearly you will think for the next fortnight.

If you raised the risk and were overruled

That is a documented organisational pattern, not a personal failure.

Information systems research has a name for the mechanism: the mum effect — the reluctance to pass on unwelcome messages, which is recognised as a significant contributor to software project failure. Bad news gets softened at each reporting layer until the version reaching the decision-makers no longer sounds like a reason to stop. Add the well-studied tendency of organisations to escalate commitment to a course of action they have already invested heavily in, and a go-live date acquires its own momentum.

This matters practically, not just emotionally. It tells you that the reporting channel was already distorted before go-live — which means it is probably still distorted now, at exactly the moment you need accurate information fastest. Assume the picture reaching you is better than reality, and go and look for yourself.

It also means the blame conversation is coming. Postpone it. There will be a legitimate lessons-learned exercise, and it will be far more useful in six weeks with evidence than in week one with adrenaline.

Three things not to do in week one

The recovery playbook

Eleven steps, grouped into four phases. The phases matter as much as the steps, because the most common recovery error is attempting week-six work in week one — rebuilding confidence while people still cannot do their jobs simply teaches them that leadership is not paying attention.

The shape follows the recovery literature. Aiyer, Rajkumar and Havelka’s staged framework for recovering troubled IS projects moves through recognition, immediate recovery, sustained recovery and maturity — and the discipline it enforces is that you do not move to the next stage until the current one holds.

Phase 1 — Stabilise (first 72 hours)

1. Triage operational harm. Before anything else, establish what is actually being damaged in the business, not in the project. Money that cannot be collected, goods that cannot ship, staff who cannot be paid, customers receiving wrong information, regulatory or statutory deadlines at risk, safety implications.

Rank by consequence and reversibility. A wrong invoice you can credit next week is not the same class of problem as a payroll error or a missed regulatory filing. Everything else — annoyance, slowness, ugly screens, unhappy users — waits. You will feel pressure to fix the loudest problem rather than the most damaging one; those are rarely the same problem.

2. Establish transparent communication. In an information vacuum people assume the worst, and they are usually right to. Publish a single, predictable update: same channel, same time each day, owned by a named person.

Say what is broken, what is being worked on now, what is scheduled next, what has been fixed since yesterday, and — the item most organisations omit — what will not be fixed soon, with the workaround to use meanwhile. Honest bad news restores more credibility than optimistic vagueness, because people can verify it against their own experience. Anything that contradicts what they see at their desk destroys trust immediately and permanently.

3. Separate technical defects from adoption problems. This is the single highest-leverage step in the playbook, and most recoveries skip it.

The daily incident list will be a mixture of at least four different things: genuine software defects, data and migration errors, process design gaps, and people not yet able to do the task. They look identical in a ticket queue and they need completely different owners, different fixes and different timelines. Treating them as one pile is how programmes end up applying a technical fix to a capability problem for six weeks.

Classify every significant issue into one of those four buckets on arrival. Where the cause is genuinely unclear, a 5 Why root cause analysis on the top few issues will usually resolve it in under an hour; where several causes are tangled together, map them with a fishbone analysis before assigning work. The classification split itself becomes management information: if 70% of your incidents are adoption problems, no amount of development capacity will fix your go-live.

Phase 2 — Protect and prioritise (weeks 1–2)

4. Protect users. Your people are absorbing the failure personally. They are being shouted at by customers and colleagues for a situation they did not create, many of them warned about it, and some are working late to compensate.

Say publicly and specifically that the problems are system and process problems, not user failure. Remove performance targets that the broken process makes unattainable. Give explicit permission to use the documented workaround rather than the “correct” route. Watch for the people quietly working three extra hours a day to hold it together, because they will break, and they are usually your best staff.

There is a hard operational reason for this beyond decency. Amy Edmondson’s research on psychological safety in work teams found that better-performing teams reported more errors, not fewer — because people were willing to speak up. In a failing go-live, your entire recovery depends on problems surfacing quickly. The moment users conclude that reporting a problem gets them blamed, your incident data becomes fiction and you lose the ability to see what is happening. Pushback in this period is information, not insubordination.

5. Prioritise critical workflows. You cannot fix everything at once, so choose openly rather than by whoever escalates loudest.

Identify the handful of end-to-end workflows the business genuinely cannot operate without — order to cash, procure to pay, payroll, statutory reporting, whatever they are in your context — and drive those to working condition before anything else. Whole workflows, not individual defects: a fixed screen in the middle of a broken chain delivers nothing. Your change impact and mitigation register already lists which roles depend on what, which makes it the fastest way to see what a given broken workflow is actually blocking downstream.

Then publish the priority order. People tolerate waiting far better when they can see where they are in the queue and why.

6. Create recovery metrics. Crises run on anecdote unless you install measurement deliberately, and anecdote is what lets a recovery drift for months without anyone being able to say whether it is working.

Keep the set small enough to publish daily on one page:

The direction of travel matters more than the absolute numbers. A steadily falling incident count with rising unaided completion is a recovery, even while the totals still look bad. Flat numbers after three weeks of effort mean the diagnosis is wrong — and these are the same principles that separate measures that drive decisions from reporting theatre.

Phase 3 — Recover (weeks 2–8)

7. Equip managers. In a bad go-live, line managers are simultaneously the most important and the least supported group. Their teams are struggling, they have no more information than anyone else, and they are being asked to maintain output.

Brief them separately and earlier than everyone else, even by only a few hours. Give them explicit authority to reprioritise their team’s work, a direct escalation route that bypasses the ticket queue for genuine blockers, and clear language for the questions they are being asked — including permission to say “this is broken and here is what is being done,” which most managers will not say unless told they may. Whether the recovery holds locally depends heavily on them, which is the same reason middle management makes or breaks go-live in the first place.

8. Log workarounds. Under this much pressure people will invent workarounds within days, and those workarounds will be numerous, undocumented and load-bearing.

Get them into a register immediately: what it is, who uses it, what it exists to bypass, whether it is safe, whether it is temporary or structural, who owns removing it, and by when. Two reasons this cannot wait. First, an unrecorded workaround becomes permanent infrastructure — this is precisely how organisations end up with spreadsheets quietly running core ERP processes years later. Second, some crisis workarounds carry real control, audit or data-integrity risk, and you need to know which ones. Anything that cannot be closed during recovery belongs in the transformation RAID log with a named owner.

9. Maintain sponsor accountability. Sponsors have a predictable failure mode after a bad go-live: they disengage. The launch was the milestone, it went badly, attention becomes uncomfortable, and the recovery gets delegated downwards to the people with the least authority to fix structural problems.

Hold the line on three things: a standing recovery review with the sponsor present, not a delegate; decisions that only they can make — funding, resource release, deprioritising other initiatives, changing decision rights — taken on a stated timetable; and visible sponsor presence with affected teams. Use stakeholder analysis to work out who else has to move for a given blocker to clear, because in recovery the binding constraint is usually a decision nobody has taken rather than work nobody has done.

Phase 4 — Rebuild and exit (week 6 onwards)

10. Rebuild confidence. Confidence does not return because the incident count fell. It returns when people experience the system working repeatedly, and when they see that raising a problem produced a fix.

So make the causal link visible: publish what was raised and what changed as a result. Close the loop by name where you can. Give people supervised repetition on the workflows they lost confidence in, because competence and confidence recover together and neither returns from reassurance alone — this is the same reason practice beats training. And use local credibility rather than programme credibility: at this point a respected peer saying “it works now, I have been using it” carries more weight than any leadership message, which is exactly what a change champion network is for.

Expect this to take longer than the technical fixes. Trust is slower to rebuild than software.

11. Decide when hypercare becomes sustainment. Hypercare usually ends by budget exhaustion rather than by decision, which leaves the organisation unsupported at precisely the wrong moment.

Set exit criteria explicitly, and check them before standing the team down:

That last cluster is the real handover. Without named business owners, the unresolved items do not transition — they simply stop being tracked, which is the mechanism by which benefits quietly leak after go-live. A recovered go-live that hands over badly still loses most of its business case.

The playbook at a glance

StepWhat good looks likeFailure mode
1. Triage operational harmRanked by business consequence and reversibilityFixing the loudest problem, not the costliest
2. Transparent communicationSame channel, same time, includes what will not be fixedOptimism that contradicts what users see
3. Separate defects from adoptionEvery issue classified into one of four causesOne undifferentiated ticket pile
4. Protect usersTargets suspended, workarounds sanctioned, blame refusedUsers absorb the failure and stop reporting
5. Prioritise critical workflowsWhole end-to-end flows fixed, order publishedIsolated defect fixes in a broken chain
6. Recovery metricsOne page, daily, trend over absolute valuesRecovery run on anecdote
7. Equip managersBriefed first, given authority and an escalation routeManagers as uninformed message-carriers
8. Log workaroundsRegistered with owner, risk and closure dateCrisis workarounds become permanent
9. Sponsor accountabilitySponsor present; structural decisions on a timetableRecovery delegated to people who cannot fix it
10. Rebuild confidenceFixes visibly traced to what users raisedDeclaring success before people feel it
11. Hypercare to sustainmentExit criteria met, business owners namedSupport ends when the budget does

Why an outside view usually helps here

There is a specific, well-evidenced reason to bring in someone external at this point, and it is not capability.

Keil and Robey’s study of turning around troubled software projects found that de-escalation — the point where an organisation stops pouring resources into a failing course of action and redirects — was in most cases triggered by actors such as senior managers, internal auditors or external consultants. In other words, the people closest to the project were rarely the ones who could call it, not because they lacked insight but because they were inside the commitment.

If you have been in the programme since design, you are subject to the same dynamic. That is worth knowing about yourself. An independent read is most valuable in the first two weeks, when the classification in step 3 is being set and the wrong diagnosis is cheapest to correct.

A structured review such as the Readiness Diagnostic Sprint exists for exactly this: separating technical from adoption causes, ranking what is actually harming the business, and producing a prioritised plan with named owners — quickly enough to matter.

Priority intervention Go-live in trouble right now? Book a 20-minute scoping call for a priority discussion — separating defects from adoption problems, ranking real business harm, and agreeing the first two weeks.
Book a 20-minute scoping call

What to do in the next hour

If you are in the middle of this today, the whole playbook is too much. Start here.

That is a day’s work and it converts a crisis into something with a shape.

A bad go-live is recoverable

Most of them are, and a fair number of systems people now rely on had a first fortnight nobody wants to discuss.

What separates the recoveries from the long, grinding failures is rarely technical skill. It is whether the organisation triaged by real harm instead of noise, told the truth early enough to keep credibility, distinguished broken software from unready people, and protected the users well enough that they kept reporting problems.

And it is whether someone decided when the crisis ended, rather than letting the budget decide.

The diagnostic tools referenced above — impact register, RAID, 5 Why and fishbone — are collected in the AI-assisted diagnostic toolkit, and the adoption risk resources cover the signals that predict this situation before it happens.

Ritvars Mētra

Ritvars Mētra

Founder of ReadinessCompass

Ritvars Mētra is the founder of ReadinessCompass, where he develops practical tools for understanding and managing organisational change complexity. His work focuses on adoption readiness, stakeholder analysis, and evidence-based change management for large-scale software and AI implementations.

View full profile →