The plan said four weeks of parallel running. Finance would key the critical transactions into both systems, reconcile daily, and switch the legacy system off at the end of the first month.
Fourteen months later the legacy system is still on. Two people are still keying into both. Nobody has formally decided to continue — there simply has never been a month where switching it off felt safe, and the cost of keeping it on is spread across a licence renewal, a hosting line and two people’s time, so it has never appeared as a decision.
This is the most common failure mode of parallel running, and it is not really a failure of the technique. It is a failure to decide, in advance, what would end it.
Three things that get called the same thing
Precision matters here, because these options carry very different costs and are routinely conflated in the same steering committee sentence.
- Parallel running. Both systems process the same real work, and outputs are compared. This is the expensive one: it doubles the transactional workload for the people doing it.
- Reversion. The new system is live; the old one is kept available, unused, so you could go back. Cheap to hold, and the option most programmes think they have. The Gateway assurance process expects feasible and tested business contingency, continuity and reversion arrangements — and the word doing the work in that phrase is “tested”. An untested reversion plan is a belief, not a control.
- Staged rollout. Different populations move at different times. The legacy system stays on because some of the business is still legitimately using it, not as a safety net.
Only the first requires anyone to do their job twice. That distinction is the whole subject of this article, because it is a people cost, and people costs are the ones that do not appear in the business case.
The case for it is real
Parallel running earns its place in a narrow set of circumstances, and it is worth stating them fairly before arguing with it.
It is the only technique that validates the new system against real production data at real volume, with a known-correct answer to compare against. Testing proves the system does what was specified. Parallel running proves it produces the same result as the thing the business currently trusts — which is a different and stronger claim. Where the output is a regulated financial figure, a payroll payment or a customer bill, that difference is worth a great deal.
It also converts an irreversible decision into a reversible one for a period, which changes the character of the go-live risk considerably.
The cost that never makes the business case
The cost usually presented is infrastructure: two sets of licences, two hosting bills, two support contracts. That number is real, quantifiable and comparatively small.
The number that matters is the operational one. Parallel running asks a team to do its job twice, and it asks this in the specific weeks when that team is slowest, least confident and most stretched — the period the enterprise systems literature calls shakedown, and which Markus and Tanis describe as the phase in which the errors of every prior phase are felt, as reduced productivity or business disruption.
So the arithmetic is not “double the work”. A team operating at perhaps 60 per cent of its normal speed in the new system, while also maintaining the old one at full speed, is being asked for something closer to 160 per cent of its previous output. For weeks. That gap is filled with overtime, with corners cut somewhere less visible, or with the reconciliation quietly not being done.
The third is the most common, and it is the one that destroys the entire value of the exercise. Parallel running without reconciliation is not a control. It is two data entry jobs.
The safety net can prevent the learning
This is the argument that rarely gets made, and it is the important one.
Markus and Tanis list the characteristic problems of the shakedown phase. Three of them are worth reading against a parallel running plan:
- maintenance of old procedures or manual workarounds
- no growth of end-user skills after initial training
- excessive dependence on a small number of key users
Parallel running does not merely coexist with all three. It institutionalises them, funds them, and puts them in the project plan.
If the old system is still authoritative, the new one is optional — and under time pressure people complete the work in whichever system they are faster in, which for several weeks is the old one. The new system becomes a thing they populate afterwards rather than a thing they work in. Skill in it does not grow, because nobody is ever forced to resolve a problem inside it; they resolve it in the legacy system and re-key the answer.
Markus and Tanis also observe that operational staff adopt workarounds to cope with early problems and then fail to abandon them once the problems are resolved. A parallel run is a workaround with executive sponsorship. It is the same mechanism that keeps spreadsheets alive after go-live, except formally approved and budgeted.
The practical consequence is that the shakedown clock does not start when the new system goes live. It starts when the old one goes off. A fourteen-month parallel run has not de-risked the transition; it has deferred it, while paying for both.
What it protects, and what it costs
| Protects against | At the cost of |
|---|---|
| Calculation and configuration errors reaching customers or regulators | Roughly 160% of normal workload during the weeks of lowest capability |
| Data migration defects that testing missed at production volume | Capability growth in the new system, which stalls while the old one is authoritative |
| An irreversible go-live decision | A soft exit date that slips indefinitely without anyone deciding |
| Loss of a known-good comparison point | Reconciliation effort that is the first thing dropped when the team is stretched |
When it is genuinely the right call
Four conditions, and the honest test is whether you can say yes to most of them:
- The output has an externally verifiable correct answer. Payroll, billing, statutory reporting, interest calculation. If the comparison is subjective, there is nothing to reconcile.
- Being wrong is materially worse than being slow. A regulator, a customer payment or a safety consequence justifies the capacity cost. Internal reporting usually does not.
- The volume is small enough to double. Two hundred payroll records, yes. Forty thousand daily order lines, no — and pretending otherwise produces the unreconciled version.
- You can staff it without taking the capacity from the people learning the new system. If the same two people are doing both, you have not built a control; you have built a bottleneck, and the expert overload that follows is entirely predictable.
Where the answer is no, the useful move is usually to narrow rather than abandon: parallel-run one process rather than all of them, or for one cycle rather than continuously, or on a sample rather than the full population.
Decide the exit before you start
A parallel run with no exit criterion will not end, because the question “is it safe to switch off yet?” has no natural answer and asking it carries all the risk while switching off carries none of the credit.
Write the exit down before go-live, in the same document as the go/no-go criteria, and make it specific:
- A match standard. Not “when we are confident” but, for example, three consecutive cycles with variances under a stated threshold, all explained.
- A maximum duration at which the decision escalates automatically, whether or not the standard is met. This is the clause that prevents month fourteen.
- A named decision-maker, and it should not be the person doing the reconciliation. They will never feel ready.
- A funded decommissioning plan. The National Audit Office is blunt on this point, recommending that bodies produce plans for the legacy estate so that maintenance, support and decommissioning are systematically addressed and the required funding is ringfenced. Ringfenced, because decommissioning budget is the first thing reallocated when a programme is over-spent, and an unfunded switch-off does not happen.
The same NAO report makes a wider point worth carrying into the design: it recommends developing interim as well as target operating models. A parallel run is an interim operating model. It has its own roles, its own workload, its own controls and its own failure modes, and it deserves to be designed rather than treated as a gap between the old world and the new one.
The cheaper things that buy most of the same safety
Stage the delivery instead. When the Bank of England renewed the Real-Time Gross Settlement system, it managed the risk of a new technology structure partly by planning delivery as a series of stages, each capable of operation in its own right. Each stage stands alone, so there is a defensible stopping point without anyone running two systems in parallel.
Keep the legacy system read-only. Most of the reassurance people want from parallel running is the ability to look something up, not to reprocess it. Read-only access delivers that at a fraction of the cost and creates no dual-entry burden.
Reconcile without re-keying. Where the legacy system can still produce its own output from migrated data, compare the outputs without asking people to enter transactions twice. The control survives; the workload does not.
Rehearse instead of hedging. A large part of the appetite for parallel running is unstated doubt about whether people can operate the new system. That is a capability question, and a rehearsal answers it directly and far more cheaply — before go-live, when there is still time to act on the answer.
The decision underneath the decision
Parallel running is often chosen not because the analysis supports it but because nobody wants to be the person who said the safety net was unnecessary. It is a decision made under governance pressure, and it is rarely revisited once made.
The NAO’s summary of why digital business change programmes fail applies here almost word for word: when they run into difficulty the technology solution is often cast as the primary reason, but these programmes face intrinsic business challenges as well as technical ones. Parallel running is a technical answer to what is usually a capability worry — and technical answers to capability worries tend to postpone the problem while consuming the budget that would have solved it.
Keep the legacy system on where being wrong is genuinely worse than being slow. Keep it on with a written exit standard, a maximum duration, a named decision-maker and ringfenced money to switch it off. And be honest that every week it stays on is a week in which the adoption you are trying to measure has not really started.
Whatever the exit standard, the switch-off itself deserves planning rather than drift. Decommissioning is the one adoption test that cannot be gamed, and it is also the one nobody is funded or measured to complete.
More on carrying and controlling go-live risk in the Adoption Risk Hub, or see how a Readiness Diagnostic Sprint tests the capability question directly, before it becomes a licensing decision.
