For a lot of companies, having a continuity plan — or a disaster recovery plan — already feels like the problem is solved. The document exists, it's saved somewhere, someone signed off on it. NIST's contingency planning guide treats that document as only half the job: the other half is testing, training, and exercising the plan — and that's the part most companies never get to.

A 2025 survey by Databarracks, a UK-based data recovery company, shows how much progress has been made in recent years: nine in ten organizations tested some part of their own recovery capability in the twelve months before the survey, a big jump from previous years. But the same survey points to an uncomfortable detail — confidence in that recovery actually dropped slightly, and the researchers' own reading is that the reason is exactly that: testing more is revealing gaps nobody used to see.

That gap is what decides how much a real outage ends up costing: the difference between keeping a plan in a folder and actually knowing, because it has been rehearsed, that it works with the people who would really have to run it.

The document nobody opens again

Testing "some part" of recovery — as the Databarracks number shows — usually means restoring one isolated server, on its own, in a controlled setting. That's useful, but it's different from gathering the people who would actually be on duty during a real outage and running the whole script: who calls whom, who decides what, who has access to which system.

That full script is what stays untouched. The plan was written once — often by a consultant, to close out an audit or satisfy a bigger client's requirement — and then filed away on a shared drive. Nobody opens it again until the day it's needed, and that's the day the company also discovers: the supplier's phone number changed, the emergency console password isn't the one written down anymore, the person responsible for calling the bank no longer works there.

ISO 22301, the international standard for business continuity management, treats that rehearsal as a mandatory step in its own management cycle, not an extra: after planning, the standard requires testing the plan and recording the result, before anything can be considered ready. A document that has never gone through that step is, by the standard's own logic, incomplete — even if it looks complete on paper.

A common format for that rehearsal, used by CISA — the US government's cybersecurity agency — as well, is the tabletop exercise: a room-based simulation, with nothing actually switched off, where each person answers live what they would do in a given crisis scenario. The agency keeps ready-made packages with more than a hundred different scenarios, precisely because most of the value of the exercise is discovering, in a calm room, what the document failed to anticipate — before discovering it in the middle of a real outage.

Why the usual approach doesn't work

The usual approach is to write the plan, file it away, and hope. The annual review, when it happens, is usually a short meeting where someone reads the document out loud and everyone agrees it looks fine — without simulating anything, without testing a single contact, without timing how long each step actually takes.

Another version of the same problem is handing out roles on paper without checking whether they still make sense: the person listed as the decision-maker may have changed jobs, the supplier listed may have closed the account, the remote-access system named in the document may have been replaced two years ago. None of that shows up from re-reading the text — it only shows up when someone actually tries to follow the script.

The most expensive effect of that habit shows up at exactly the wrong moment: in the middle of the actual outage, when the team discovers live that a step in the plan depends on something that no longer exists. At that point nobody is testing anything — they're improvising under pressure, while the business stays down.

Confidence in recovery dropped slightly — a sign that testing is revealing gaps nobody used to see.

Data Health Check 2025, Databarracks

What has to be in place

A continuity plan that actually works when it's needed rests on a few verifiable mechanisms — not on good intentions.

An exercise on the calendar, actually run, not left for whenever there's spare time. Without a fixed date, the rehearsal becomes the first thing cut from the schedule, month after month.

Supplier, bank, and carrier contacts checked again every cycle, not copied from the last document without confirming they still work.

Emergency passwords and access kept somewhere that works even without the one person who usually logs in. If only one person knows how to get into the emergency console, the plan depends on that person being available — and they might be on vacation on the exact day of the outage.

A specific role for whoever is on duty that day, not just for whoever wrote the plan. Whoever is on duty is rarely the person who drafted the document.

A record of what the exercise revealed, with a deadline to fix it before the next cycle — without that record, the same problem shows up again at the next rehearsal.

The full script timed from start to finish, to know how long it actually takes — not how long it looks like it takes on paper.

This is how Skills IT works: with the continuity exercise scheduled and logged, and supplier contacts and emergency access reviewed every cycle instead of copied from the last document.

What changes once the plan has been rehearsed

The gain from rehearsing the plan isn't abstract. In a real outage, the difference between a team that has already been through the script and one reading the document for the first time shows up in minutes: the first knows who calls, who decides, and where the access is; the second spends that time looking for it.

That changes the kind of conversation that happens during the crisis. Instead of deciding, live and under pressure, who is responsible for calling the client or when to give up on a fix and move to the backup plan, the team already knows the answer — because it already answered that same question, on a calm day, during the exercise.

The avoided cost shows up in three places: less staff time spent reconstructing decisions that should have already been made, less risk of a supplier or client losing confidence over a delayed notice, and less chance of the emergency budget going toward a problem the rehearsal itself would have already caught.

Questions to bring to the next meeting

Before reviewing the plan again just by reading the text, it's worth gathering whoever would be on duty and asking:

  1. Has anyone actually followed this plan from start to finish, with the real people involved, outside of an actual outage? If the answer is no, the plan has never been tested — only written.
  2. Do the phone numbers and contacts named in the document still work? Calling to check, today, takes only a few minutes.
  3. Who fills in for the person who always handles this, if they're on vacation or unreachable on the day of the outage? A plan with a single point of responsibility isn't a plan — it's a dependency.
  4. How long does the script actually take, timed? The answer is rarely the same as it looks on paper.
  5. What did the last exercise reveal — and has it already been fixed? A problem found and left unfixed will show up again, only next time it might be during a real outage.