How Cloud Resilience and a Disaster Recovery Plan Differ
A backup answers an important question: Do we have a copy of our data? A good disaster recovery plan asks a different one: How quickly can we restore the services the business needs to operate?
Consider a critical application that has been backed up and successfully restored during an annual test. That does not necessarily mean the business can resume operations quickly after an outage or ransomware attack. The application may depend on identity services, storage, networking or other components that also need to be restored and available in the right sequence.
That’s why recovery planning needs to start with the business impact. IT leaders should establish recovery time objectives (RTOs) — how quickly a workload needs to be operational — as well as recovery point objectives (RPOs), which determine how much data loss is acceptable.
Not every workload needs to be recovered at the same speed. The goal isn’t necessarily to bring the entire production environment back online simultaneously. Instead, organizations should identify the applications that are most important to restoring functional business operations —point-of-sale systems, for example — and prioritize them accordingly. The cloud can make that recovery more achievable through capabilities such as workload replication and failover. But those capabilities still need to be incorporated into a deliberate recovery strategy.
READ MORE: Why should small businesses modernize their IAM programs?
Build and Test a Disaster Recovery Plan That Works
One of the biggest gaps in disaster recovery is the assumption that a plan works simply because it exists. Organizations may have a recovery document that was created a year or two ago, but applications, infrastructure, security requirements and dependencies can change.
Testing exposes those gaps. A formal recovery strategy should map application dependencies, establish the infrastructure and networking required for failover, and document the procedures in a recovery runbook. Organizations can then conduct failover tests to determine whether applications meet their RTOs.
Frequent testing doesn’t necessarily have to disrupt production. Cloud-based recovery capabilities can allow organizations to test recovery processes without taking production systems offline. That makes it easier to continually validate that the recovery plan still works as the environment evolves.
The process can also identify opportunities for automation. For a particularly critical application, for example, an organization may decide that portions of the recovery runbook should be automated rather than relying on someone to execute a series of manual steps during an already stressful incident.
For organizations that already have backups and a disaster recovery plan, this process can reveal the unknowns they may have overlooked. For those without a formal plan, it provides a structured way to establish recovery fundamentals based on business requirements.
If you’ve been calming your nerves with reassuring but dubious assumptions, it’s time to replace those assumptions with facts. Having your data in the cloud doesn’t mean your business is automatically prepared for disaster. Resilience comes from knowing what must be recovered, how quickly it must be recovered, what it depends on and — critically — proving through regular testing that the plan can actually get the business back online.

