Onam

What is the difference between RTO and RPO?

In short

Recovery Time Objective (RTO) is the longest a system can be unavailable after a disruption before the impact becomes unacceptable. Recovery Point Objective (RPO) is the most data, measured in time, that can be lost — how far back the restored copy may be. RTO is about downtime; RPO is about data loss.

RTO and RPO in one example

An order database fails at 14:00. The last usable copy is from 13:45, and the service is back at 15:30.

  • The recovery point was 13:45 — fifteen minutes of orders were lost. If the RPO was 30 minutes, it was met.
  • The recovery time was 90 minutes. If the RTO was one hour, it was missed.

The two objectives are independent. A system can lose no data and still be down for a day, or come back in minutes with an hour of data missing.

How are RTO and RPO set?

They are business decisions, not technical ones. A business impact analysis asks, for each application, what an hour of downtime costs and what losing an hour of data costs — in revenue, regulatory exposure, customer harm and reputation. The answers set the targets.

Targets usually come in tiers: a handful of critical applications with tight objectives, more with moderate ones, and the rest recoverable within a day or longer. Tiering matters because tighter objectives cost more to meet.

What drives RPO?

RPO is set by how often a copy is made and how far behind it runs:

  • Periodic backups and snapshots — the worst-case data loss is roughly the interval between them, plus the time to complete one.
  • Asynchronous replication — data loss is the replication lag at the moment of failure, which varies with load.
  • Synchronous replication — writes are acknowledged only once both copies have them, so data loss can approach zero, at the cost of latency and distance limits.

What drives RTO?

RTO is set by everything that has to happen before users are served again:

  1. Detecting the failure and deciding to recover.
  2. Provisioning or activating infrastructure in the recovery location.
  3. Restoring or promoting data.
  4. Starting services in dependency order — the database before the application, the application before the load balancer.
  5. Redirecting traffic, including DNS changes and their time to live.
  6. Verifying that the application actually works.

Steps that can run in parallel reduce the total; the longest chain of steps that must run one after another — the critical path — sets it.

Required, predicted and actual

A useful discipline is to keep three numbers apart for each application:

NumberWhere it comes from
RequiredThe target the business set
PredictedWhat the current design and protection should achieve
ActualWhat a test or real recovery measured

Plans fail when these are blurred — when the required number is quoted as if it had been measured, or a prediction is never checked against a test.

RTO and RPO vs MTD

Maximum tolerable downtime (MTD) — also called maximum tolerable period of disruption — is the point beyond which the business suffers unacceptable or irreversible harm. RTO must sit inside it, with enough margin to verify the recovery and resume normal operation.

Next steps

Frequently asked questions

What does RTO stand for?

RTO stands for Recovery Time Objective: the maximum acceptable time between a disruption and the restoration of a service. It is a target set by the business, usually through a business impact analysis.

What does RPO stand for?

RPO stands for Recovery Point Objective: the maximum acceptable amount of data loss, expressed as time. An RPO of 15 minutes means the recovered data may be at most 15 minutes older than the moment of failure.

Which is more important, RTO or RPO?

Neither in general — it depends on the application. A payments ledger may tolerate some downtime but no data loss, so its RPO is tight. A public website may tolerate losing recent content but not long outages, so its RTO is tight. Each application needs both set deliberately.

Can RPO be zero?

Close to it, with synchronous replication, where a write is only confirmed once every copy has it. That adds latency and limits how far apart the copies can be, so true zero RPO is usually reserved for the most critical data. It also does not protect against corruption or deletion, which replicate too.

How do you know an RTO can actually be met?

By testing. A design and a plan give a predicted recovery time; only a drill or a real recovery measures the actual one. Recording the prediction before each test, and comparing it with the measured result afterwards, is what shows whether the plan is realistic.

See it on your own cloud

Onam DRM maps applications and dependencies, reads configured backup and replication, and predicts RTO and RPO against your targets.