Ask a programme what its day-one support model is, six weeks out, and the answer is usually some version of this: there is a mailbox, the super users will help their teams, and the project team will be around for the first two weeks.
None of that is a model. It is a hope, expressed as a staffing plan, and it fails in a predictable order: the mailbox floods on the first morning, the super users are pulled back into their own backlog by lunchtime on day two, and by the end of the first week the entire support capability of the organisation is one person who happens to know the answer and cannot leave their desk.
What day one actually produces
The first misconception is about the nature of the demand. Programmes plan support around defects, because defects are what testing taught them to expect. Day one does not mainly generate defects. It generates four quite different things:
- “How do I…” — navigational questions from people who were trained six weeks ago. High volume, low complexity, and answerable by a peer.
- “Is this right?” — confidence checks. Someone has done the task correctly and does not trust the result. Low complexity, but only a knowledgeable person can answer it, and it is the category most often mistaken for a defect.
- “This will not let me…” — genuine blocks. Authorisations, missing master data, configuration that does not fit the real case. Medium volume, and the category that actually stops work.
- “The system is wrong” — actual defects. The smallest category by a wide margin, and the only one the support plan was designed for.
A model built for the fourth category will be overwhelmed by the first two, and it will misroute the third — sending authorisation problems to a technical queue where they sit for two days while somebody cannot dispatch an order.
Borrow the staffing arithmetic from people who have solved it
Two disciplines have thought harder about surge support staffing than change management has, and both publish their numbers.
Span of control, from incident management. The US Incident Command System, taught by FEMA and used for everything from wildfires to mass-casualty events, holds that the optimal span of control is one supervisor to five subordinates, with an effective range of three to seven. Below three is inefficient; above seven, one person can no longer manage the group during an incident.
Now compare that with a typical go-live, where one floorwalker is assigned to a department of thirty people on their first day using an unfamiliar system. That is not support. At 1:30 the floorwalker is a queue, and the people at the back of it stop asking and start guessing.
Sustainable load, from site reliability engineering. Google’s SRE practice sets an explicit ceiling: a maximum of two events per 8–12 hour on-call shift. The reasoning is not comfort. It is that beyond that volume, incidents cannot be handled accurately, cleaned up properly, or learned from — the responder is reduced to firefighting, and the organisation stops improving. Google also caps aggregate operational work at 50 per cent of an SRE’s time, precisely so that operational load does not consume the capacity that makes the service better.
The translation to hypercare is direct. A support person handling forty queries a day is not diagnosing anything. They are dispatching, and nothing they learn is being captured, which is why week three looks exactly like week one.
The four tiers, and who belongs in each
| Tier | Handles | Staffed by | Ratio |
|---|---|---|---|
| 0 — At the desk | “How do I” and “is this right” | Local super users, released from their day job in writing | 1 per 5–8 users, first two weeks |
| 1 — Process help | Blocks, authorisations, master data, exceptions | Process owners and business analysts who know the design decisions | 1 per function, on a rota |
| 2 — Technical | Genuine defects, integrations, performance | The delivery team and the integrator | Per severity, with named on-call |
| Command | Prioritisation, trade-offs, the decision to stop processing | One operational leader with actual authority | One, always identifiable |
Tier 0 is the one that gets cut, and it is the one carrying eighty per cent of the volume. Cutting it does not save the cost; it moves that volume into tier 1 and tier 2, where each question costs several times as much to answer and takes hours instead of seconds.
Released in writing, or not released at all
The single most common design failure is treating super user support as something people do alongside their normal work.
It never happens. Their own queue is measured, their support role is not, and their manager is accountable for the first and not the second. By day three they are back on their own work and answering questions in the gaps, which means questions queue behind their backlog.
The fix is unglamorous and has to happen weeks before go-live: an explicit, written release of named individuals for a defined percentage of their time, agreed with the manager who owns their targets, with the backfill or the reduced target stated. If a manager will not agree it, that is not an administrative problem — it is a finding about whether the business is actually ready to absorb this change, and it belongs in front of the steering committee.
Markus and Tanis identify excessive dependence on a small number of knowledgeable key users as one of the characteristic pathologies of the post-go-live period — the organisation leans on individuals instead of building capability across all operational staff. An unstaffed support model guarantees it: when there is no tier 0, every question routes to whoever is known to have the answer, and that person becomes a bottleneck who cannot take leave for a quarter. Expert overload is not an accident of go-live; it is what happens when the support model is left to self-organise.
Three questions that break most models
Who can decide to stop processing? On day two, someone will discover that a batch has been posting incorrectly for six hours. The decision to halt is commercial, not technical, and it usually needs to be made in minutes. If there is no named person with that authority, the decision defaults to whoever is most senior in the room, or worse, gets deferred until the next call. Name them, tell everyone who they are, and give them a deputy.
What happens at 17:00 on Friday? Most support models are designed for the working week and go-lives frequently land near a period end. Publish the hours, publish who is on call outside them, and be honest if the answer is nobody — because an unpublished gap is discovered by a user at 18:30 with a customer waiting.
Where do answers go? The same question asked by five people is a documentation defect, not five support incidents. Without a mechanism for capturing an answer once and publishing it, the support team re-answers the same twenty questions for three weeks. A shared page updated twice daily is enough; the discipline is what matters, not the tooling.
This is a handover, not an extension of the project
The Gateway assurance regime treats support arrangements as a precondition of going live rather than a follow-on activity. Its Gate 4: Readiness for service review expects, as evidence that an organisation is ready for business change, a clearly defined service management function already in place — and it explicitly confirms that arrangements exist for handover of the project from the project owner to the operational business owner, with ownership clearly identified after that handover.
That last clause is the one worth arguing about internally. On most programmes, day-one support is staffed by the project, which means the people answering questions are the people who are also being disbanded. The knowledge leaves with them. If the receiving service function is not in the room from day one — ideally answering tier 2 alongside the project team — the handover is a document, and the organisation will rediscover every lesson in month four.
Measure the right thing, and know when to stop
Falling ticket volume is the metric every hypercare report leads with, and on its own it is close to meaningless. Volume falls when people stop asking, which happens both when they have learned and when they have given up and invented a workaround. Those two look identical on the chart and could not be more different underneath — which is why hypercare metrics are not adoption metrics.
Four measures that carry more signal:
- Mix by category across the four types above. A healthy trajectory is “how do I” falling while “this will not let me” stays flat — that is learning. Everything falling together is usually disengagement.
- Repeat askers. The same person asking daily in week three is a training gap that no amount of support will close.
- Concentration. If 70 per cent of questions come from one function, the problem is that function’s design or capability, not the support model.
- Answer reuse rate. How many queries were resolved from the published answers rather than from scratch.
Set the exit criteria for hypercare before it starts, for the same reason you set go/no-go criteria in advance: once it is running, nobody will feel ready to stand it down, and it will either drift for months or be cancelled abruptly on a budget date. A defensible standard looks like sustained tier-0 demand below a stated level, no severity-one incidents for a defined period, and the receiving service function handling the queue unaided for two consecutive weeks — the same “unaided” test that makes a rehearsal worth running in the first place.
The point of the model
A go-live support model is not a safety net for the software. The software has been tested. It is a capability bridge for people who know roughly what to do and cannot yet do it quickly, under pressure, with a customer waiting — and its job is to carry them across that gap fast enough that they never need to invent a workaround.
Every workaround invented in the first fortnight because nobody was available to answer a thirty-second question will still be in use next year. That is the real return on tier 0, and it is why the cheapest tier is the one worth protecting when the budget conversation starts.
Guidance tooling can absorb a large share of the first category. It is worth being clear about what it reaches: a digital adoption platform answers how-do-I very well and cannot touch the confidence checks and blocked tasks that actually stop work.
More on capability and capacity around go-live in the Adoption Risk Hub, or see how a Readiness Diagnostic Sprint sizes support demand from evidence rather than from optimism.
