The recovery plan lives in a document that was reviewed for last year's audit.
Since then a team added a database without a replica, another moved a service to a new region, and the runbook still lists the old order of start-up. Nobody is hiding anything; nobody is checking either. The plan says the application comes back in an hour. The only way to find out it will not is the outage itself, or a drill nobody has had time to schedule.
The cost of an untested plan
A recovery time that is wrong on the high side is a nuisance. Wrong on the low side, you find out during the outage. DRM is built to show the gaps before then.
The mechanism, not the marketing
- 1
DRM starts from the Onam platform's cloud inventory rather than a separate scan, and places each resource where the provider says it lives — its availability zone where one is given, its region otherwise. It does not guess a placement the provider does not state.
- 2
Applications are proposed from your tags, each with the tag it came from and a confidence in what that tag means. An infrastructure grouping such as a Kubernetes cluster is labelled as one, and resources that could not be grouped are shown as plainly as those that could.
- 3
Dependencies come from a catalogue of how cloud resources relate, and each edge says what kind it is: something that must recover first, something that proves a copy exists, or something that shares a failure domain. Only the first kind sets recovery order.
- 4
Protection is read from cloud configuration — backup plans and retention, snapshots, replicas, multi-zone and cross-region replication. A configured mechanism is reported as found, never as proven: configuration does not show whether last night's backup succeeded or whether a restore has ever worked, and DRM does not pretend it does. There are no connectors to backup products.
- 5
Recovery plans are composed from approved applications, dependencies and protection: ordered steps, work that can run in parallel, and for every step whether its position came from a discovered dependency or from convention.
- 6
Predicted RTO is the critical path through the approved plan; predicted RPO is the worst replication link, taking the larger of observed lag and configured target. Where a figure cannot be derived — no plan, or backup-only protection with no recorded frequency — it is left blank rather than guessed. Required RTO and RPO are your inputs, and a calculated value never overwrites them.
- 7
Engines propose and people approve. Approved models are frozen as baselines, and drift is measured against the baseline you signed off — not against last week's scan — and ranked by what it does to recoverability. A rejected proposal is remembered and not proposed again.
- 8
DR drills run in your own tooling are recorded in DRM — planned, running, completed — with the measured recovery time and data loss set against the prediction captured when the drill was planned. DRM records drills; it does not run them or execute a failover.
Specific outputs, measurable outcomes
Questions we get a lot
Ready to see Onam DRM on your applications?
DRM runs in the same console as Estate, Security and FinOps, on the same inventory — your cloud accounts are connected once.