Onam

What is cloud disaster recovery?

In short

Cloud disaster recovery is the set of strategies, configurations and procedures that restore applications and data after a major disruption — a regional outage, data corruption, deletion or a cyberattack — using cloud infrastructure as the recovery site. Strategies range from restoring backups to running fully active copies in more than one region.

Disaster recovery vs high availability

High availability keeps a service running through small, expected failures: an instance dies, a disk fails, an availability zone goes dark. It is designed into the architecture and usually works automatically.

Disaster recovery handles large or unusual failures that high availability does not cover: a whole region unavailable, data deleted or encrypted by an attacker, a bad change replicated everywhere at once. Replication alone is not disaster recovery — a corrupted table replicates as faithfully as a good one. That is why point-in-time copies, isolated from the primary, remain part of every serious plan.

What are the cloud disaster recovery strategies?

Four strategies are commonly described — including in cloud providers' own disaster recovery guidance. They trade cost against recovery time and data loss:

StrategyWhat runs in the recovery regionTypical RTO / RPORelative cost
Backup and restoreNothing but copies of data and infrastructure definitionsHoursLowest
Pilot lightData replicated continuously; core services provisioned but switched off or scaled to zeroTens of minutesLow
Warm standbyA complete, scaled-down copy of the environment, runningMinutesMedium
Multi-site active/activeA full copy serving live traffic in more than one regionNear zeroHighest

The figures are orders of magnitude, not guarantees. Actual recovery time depends on the application, its data volume and the procedure — which is why it has to be tested.

Backup and restore

Data is backed up — ideally to another region and account — and infrastructure is defined as code so it can be recreated. After a disaster, the environment is rebuilt and data restored. Cheapest to run; slowest to recover.

Pilot light

The data layer is replicated continuously to the recovery region, and the rest of the environment exists in a minimal state. Recovery means switching on and scaling up the application tier around data that is already there.

Warm standby

A functional, smaller copy of the whole environment runs all the time. Recovery means scaling it to production size and redirecting traffic. Because it is always running, it can also be tested more easily.

Multi-site active/active

Two or more regions serve production traffic simultaneously. Losing one means the others absorb its load. It gives the fastest recovery but requires the application to handle data consistency across regions, which is a design decision, not a configuration switch.

How to choose a strategy

Start from the business targets, not the technology. A business impact analysis gives each application a required RTO and RPO; the strategy is the cheapest one that meets them. Most organisations use several at once — active/active for a few critical services, backup and restore for most internal systems.

What makes a disaster recovery plan work?

  1. A complete inventory of what each application depends on — the database, the queue, the secrets, the DNS records, the identity provider.
  2. A recovery order that follows those dependencies.
  3. Protection that matches the targets — backup frequency and replication that can actually deliver the RPO.
  4. Isolation — copies in a separate account or with immutability, so an attacker or a mistake cannot delete the primary and the backup together.
  5. Testing — regular drills that measure real recovery time and data loss against the targets.
  6. Drift control — a check that the environment still matches the plan, because a new database added without a replica quietly breaks it.

Regulation

Disaster recovery is increasingly a regulatory expectation, not only good practice. Financial-sector rules such as the EU's Digital Operational Resilience Act (DORA) expect tested ICT business continuity and recovery arrangements, and healthcare rules such as the HIPAA Security Rule require a contingency plan including data backup and disaster recovery procedures.

Next steps

Frequently asked questions

What is the difference between backup and disaster recovery?

A backup is a copy of data. Disaster recovery is the whole capability to bring an application back: the copies, the infrastructure to run on, the order in which components start, the procedure, and the testing that proves it works. Backups are one input to disaster recovery, not a substitute for it.

Is multi-region replication enough for disaster recovery?

Not on its own. Replication protects against losing a location, but it also copies deletions, corruption and ransomware encryption to the other region. Point-in-time backups kept isolated from the primary are still needed, along with a tested procedure for failing over.

How often should disaster recovery be tested?

At least as often as regulation or internal policy requires, and after significant changes to the application. Many organisations test critical applications several times a year and the rest annually. What matters is that each test measures actual recovery time and data loss against the targets.

What is pilot light in disaster recovery?

A strategy where the data layer is continuously replicated to the recovery region while the rest of the environment is kept minimal or switched off. In a disaster, the application tier is started and scaled up around data that is already in place.

What is drift in disaster recovery?

Drift is any change to the environment since the recovery plan was approved that the plan does not reflect — a new resource without protection, a dependency that moved, a changed recovery site. Undetected drift is how a plan that was correct when signed becomes wrong without anyone editing it.

See it on your own cloud

Onam DRM maps applications and dependencies, reads configured backup and replication, and predicts RTO and RPO against your targets.