What is cloud disaster recovery?
Cloud disaster recovery is the set of strategies, configurations and procedures that restore applications and data after a major disruption — a regional outage, data corruption, deletion or a cyberattack — using cloud infrastructure as the recovery site. Strategies range from restoring backups to running fully active copies in more than one region.
Disaster recovery vs high availability
High availability keeps a service running through small, expected failures: an instance dies, a disk fails, an availability zone goes dark. It is designed into the architecture and usually works automatically.
Disaster recovery handles large or unusual failures that high availability does not cover: a whole region unavailable, data deleted or encrypted by an attacker, a bad change replicated everywhere at once. Replication alone is not disaster recovery — a corrupted table replicates as faithfully as a good one. That is why point-in-time copies, isolated from the primary, remain part of every serious plan.
What are the cloud disaster recovery strategies?
Four strategies are commonly described — including in cloud providers' own disaster recovery guidance. They trade cost against recovery time and data loss:
| Strategy | What runs in the recovery region | Typical RTO / RPO | Relative cost |
|---|---|---|---|
| Backup and restore | Nothing but copies of data and infrastructure definitions | Hours | Lowest |
| Pilot light | Data replicated continuously; core services provisioned but switched off or scaled to zero | Tens of minutes | Low |
| Warm standby | A complete, scaled-down copy of the environment, running | Minutes | Medium |
| Multi-site active/active | A full copy serving live traffic in more than one region | Near zero | Highest |
The figures are orders of magnitude, not guarantees. Actual recovery time depends on the application, its data volume and the procedure — which is why it has to be tested.
Backup and restore
Data is backed up — ideally to another region and account — and infrastructure is defined as code so it can be recreated. After a disaster, the environment is rebuilt and data restored. Cheapest to run; slowest to recover.
Pilot light
The data layer is replicated continuously to the recovery region, and the rest of the environment exists in a minimal state. Recovery means switching on and scaling up the application tier around data that is already there.
Warm standby
A functional, smaller copy of the whole environment runs all the time. Recovery means scaling it to production size and redirecting traffic. Because it is always running, it can also be tested more easily.
Multi-site active/active
Two or more regions serve production traffic simultaneously. Losing one means the others absorb its load. It gives the fastest recovery but requires the application to handle data consistency across regions, which is a design decision, not a configuration switch.
How to choose a strategy
Start from the business targets, not the technology. A business impact analysis gives each application a required RTO and RPO; the strategy is the cheapest one that meets them. Most organisations use several at once — active/active for a few critical services, backup and restore for most internal systems.
What makes a disaster recovery plan work?
- A complete inventory of what each application depends on — the database, the queue, the secrets, the DNS records, the identity provider.
- A recovery order that follows those dependencies.
- Protection that matches the targets — backup frequency and replication that can actually deliver the RPO.
- Isolation — copies in a separate account or with immutability, so an attacker or a mistake cannot delete the primary and the backup together.
- Testing — regular drills that measure real recovery time and data loss against the targets.
- Drift control — a check that the environment still matches the plan, because a new database added without a replica quietly breaks it.
Regulation
Disaster recovery is increasingly a regulatory expectation, not only good practice. Financial-sector rules such as the EU's Digital Operational Resilience Act (DORA) expect tested ICT business continuity and recovery arrangements, and healthcare rules such as the HIPAA Security Rule require a contingency plan including data backup and disaster recovery procedures.
Next steps
- What is the difference between RTO and RPO? — the two targets every strategy is measured against
- How Onam DRM maps recovery readiness — dependencies, protection, predicted RTO/RPO and drift
- What is a cloud asset inventory? — the starting point of every recovery plan
Frequently asked questions
What is the difference between backup and disaster recovery?
A backup is a copy of data. Disaster recovery is the whole capability to bring an application back: the copies, the infrastructure to run on, the order in which components start, the procedure, and the testing that proves it works. Backups are one input to disaster recovery, not a substitute for it.
Is multi-region replication enough for disaster recovery?
Not on its own. Replication protects against losing a location, but it also copies deletions, corruption and ransomware encryption to the other region. Point-in-time backups kept isolated from the primary are still needed, along with a tested procedure for failing over.
How often should disaster recovery be tested?
At least as often as regulation or internal policy requires, and after significant changes to the application. Many organisations test critical applications several times a year and the rest annually. What matters is that each test measures actual recovery time and data loss against the targets.
What is pilot light in disaster recovery?
A strategy where the data layer is continuously replicated to the recovery region while the rest of the environment is kept minimal or switched off. In a disaster, the application tier is started and scaled up around data that is already in place.
What is drift in disaster recovery?
Drift is any change to the environment since the recovery plan was approved that the plan does not reflect — a new resource without protection, a dependency that moved, a changed recovery site. Undetected drift is how a plan that was correct when signed becomes wrong without anyone editing it.