Your continuity plan probably models the wrong disaster
Most plans are built around losing something physical: a site, a data centre, a hosting region. That scenario is easy to describe and easy to cost, so it becomes the scenario the plan is written for. The incidents that actually halt a scaling business are usually identity compromise, a supplier outage, or a corrupted integration quietly writing bad data into the ERP for a fortnight.
The distinction that matters is between an availability failure and an integrity failure. In an availability failure you know the data is sound and you only need to restore access. In an integrity failure you do not know when the estate stopped being trustworthy, so the first job is establishing a clean point in time. That investigation, not the restore, sets the recovery clock, and it is the part no runbook has a duration for.
Recovery order matters more than recovery speed
Businesses tend to buy faster restore. The binding constraint is more often sequence. Identity has to come back before anything that authenticates against it. The integration layer has to come back before the systems that exchange data through it. Customer-facing channels should come back last, because they generate obligations the rest of the estate has to be ready to honour.
The failure is concrete. Restore the ERP first while the identity provider is still down and nobody can log in. Bring the webshop up before the order integration and you accumulate orders that cannot be fulfilled, then spend days unwinding them by hand. Speed applied in the wrong order creates a second incident inside the first.
Backups fail in the ways nobody tests
Almost every organisation tests the restore of a single file or a single virtual machine. Very few have restored the whole estate, in order, under time pressure, with the people who would actually be available. A successful backup job proves that data was written. It does not prove you have a recovery capability.
Three traps recur. The backup console authenticates against the same directory the attacker now controls, so the credentials needed to recover are inside the blast radius. Hold a break-glass account and its recovery path outside the estate. Retention is shorter than dwell time, so every recoverable copy already contains the problem. And there is nowhere clean to restore onto, because the only infrastructure available is the infrastructure you no longer trust.
The honest counter-argument: much resilience spending buys the wrong protection
The conventional advice is to invest in high availability: replication, hot standby, multi-region. Those controls protect against equipment and site loss, which is the least likely thing to happen to most scaling businesses. Worse, replication is obedient: it will faithfully copy corrupted or encrypted data to the standby, at speed.
The cheaper and usually higher-return investment is an agreed degraded operating mode. What does the business actually do for seventy-two hours with no ERP? Which orders can be taken on paper, under whose authority, at what prices, and how are they reconciled afterwards? Who holds an offline copy of the open order book, the customer contact list and the supplier terms?
Decision authority is the scarce resource in the first hour
Technical containment is often quick. Decisions are slow. Someone has to authorise taking the ERP offline, disconnecting a supplier link, telling customers, engaging insurers and legal counsel, and committing unbudgeted spend. In most organisations every one of those sits with a person who is not on the incident bridge.
Pre-delegate. Name an incident lead and an alternate, give the role a standing spend authority, and agree an out-of-band communication route on the assumption that email and the usual chat tool are untrusted. Keep a decision log from the first minute: the questions asked afterwards are about what you knew and when. UK GDPR places a short notification clock on personal data breaches, so decide who assesses and who signs before you need to.
Your recovery time cannot be shorter than your suppliers’
Hosting, third-party logistics, payments, EDI, the managed service desk: their incident is your incident, and you will be managing it with no access to their systems and no control over their communications. A recovery objective that assumes suppliers respond at your pace is a number you have written for yourself.
Assess them on the right things. A certificate tells you a management system existed on the day of audit. What you need is their incident communication process, the recovery commitment in the contract, and a named human who answers on a Sunday. Ask about concentration too: one supplier behind several critical services is a single point of failure with several invoices.
Where the contract is silent on notification of incidents affecting your data, that is a renewal conversation. Cyber governance and ISO 27001 work is at its most useful here, because it puts supplier risk into a register with an owner rather than leaving it in a procurement folder.
Exercise decisions, not procedures
Most tabletop exercises walk the runbook with the right people present, complete information and working tools. That rehearses the easy part. Design the exercise to remove things: the key person is unavailable, email is down, the information is partly wrong, the supplier is not answering. Then measure how long each decision took and what it stalled on.
Feed the findings into the architecture backlog, not only into the plan document. If the exercise shows the webshop cannot run for an hour without the ERP, that is a coupling problem to be designed out, not a paragraph to be reworded. If it shows nobody can list which integrations touch customer data, that is a data-mapping task.
Resilience is a property of the architecture and of the organisation running it. An exercise is the cheapest instrument available for observing both, and the only one you can use before the day it matters. Sequencing that work alongside everything else competing for attention is what fractional CIO engagements are usually brought in to do.
Test the plan before the incident does
Link-IT works with leadership teams to turn continuity assumptions into a tested recovery sequence, clear incident authority and an architecture that degrades gracefully rather than stopping.
Discuss your next stage