0%

Preparing the page

Cloud Disaster Recovery: RTO, RPO and Choosing the Right DR Strategy

A practical guide to cloud disaster recovery: RTO and RPO explained, four DR strategies from backup-and-restore to active-active, costs and how to test DR.

Nidhish Joy

Nidhish Joy · Co-founder & CEO

· 4 min read

Share
Placeholder illustration

"We are in the cloud, so we are resilient." It is a common belief and a dangerous one. Cloud providers offer the building blocks for resilience, but designing, implementing and testing disaster recovery is still your responsibility.

This guide explains cloud disaster recovery in practical terms.

Cloud providers offer the building blocks for resilience, but designing, implementing and testing disaster recovery is still your responsibility.

What are RTO and RPO?

  • Recovery Time Objective (RTO): the maximum acceptable time to restore a service after a disaster. "Checkout must be back within 1 hour."
  • Recovery Point Objective (RPO): the maximum acceptable data loss, measured in time. "We can lose at most 5 minutes of orders."

Set RTO and RPO per application, based on business impact, not one number for everything.

RTO and RPO timeline

What disasters should cloud DR cover?

  • Regional or availability zone outages
  • Ransomware and malicious deletion
  • Accidental deletion or corruption by people or automation
  • Failed deployments or migrations
  • Compromised accounts
  • Third-party or SaaS dependency failures

What are the four main DR strategies?

StrategyHow it worksTypical RTOTypical RPORelative cost
Backup and restoreBack up data; rebuild infrastructure when neededHours to a dayHoursLow
Pilot lightCore data replicated; minimal infrastructure ready to scaleTens of minutes to hoursMinutesLow–medium
Warm standbyScaled-down full copy always runningMinutesSeconds to minutesMedium–high
Multi-site active-activeFull capacity in multiple regionsNear zeroNear zeroHigh

These strategies are described in AWS's disaster recovery guidance and apply similarly on Azure and Google Cloud.

Four DR strategies by cost and RTO

How do you choose the right strategy?

  1. Classify applications into tiers by business impact (for example, Tier 1 revenue-critical, Tier 2 important, Tier 3 internal).
  2. Set RTO and RPO per tier with business owners.
  3. Map each tier to a strategy that meets the target at acceptable cost.
  4. Consider dependencies: identity, DNS, networking and shared databases must also recover.
  5. Automate recovery with infrastructure as code so environments can be rebuilt reliably.

How do you protect against ransomware?

Traditional replication can copy encrypted or deleted data to your DR site. Add:

  • Immutable backups (for example, object lock or vault lock)
  • Separate backup accounts with restricted access
  • Point-in-time recovery for databases
  • Regular restore tests from immutable copies

See our cloud backup strategy guide.

How often should you test DR?

Test at least annually for every tier, and more often for Tier 1 systems. Test types include:

  • Tabletop exercises: walk through the runbook as a team
  • Component tests: restore a database or rebuild an environment
  • Full failover tests: move production to the DR site and back

Measure actual recovery time and data loss against RTO and RPO, and fix gaps.

How Crozaint approaches disaster recovery

Disaster Recovery and Business Continuity is one of the eight disciplines in Crozaint's 24/7 managed services, alongside storage and backup, database services and cloud infrastructure operations. We help classify applications, design DR to meet RTO and RPO targets across AWS, Azure and Google Cloud, automate recovery with tools like Terraform and Ansible, and run regular DR tests.

Because we also operate your environment day to day under ITIL-based processes, DR runbooks stay current as systems change, rather than going stale in a document.

Common mistakes to avoid

  • One RTO and RPO for every system
  • Replicating ransomware-encrypted data to DR
  • Forgetting identity, DNS and network dependencies
  • Never testing failover
  • DR runbooks that do not match the current architecture

Conclusion

Cloud DR is a business decision implemented with technology. Set RTO and RPO per application, choose a fitting strategy, protect against ransomware and test regularly.

Not sure your DR plan would work today? Book a 30-minute call with Crozaint.

Frequently Asked Questions

What is the difference between RTO and RPO?

RTO, recovery time objective, is how quickly a system must be restored after a disaster. RPO, recovery point objective, is how much data loss is acceptable, measured in time. For example, an RTO of one hour and RPO of five minutes means recovery within an hour, losing at most five minutes of data.

Is the cloud automatically disaster-proof?

No. Cloud providers offer highly resilient infrastructure and tools such as multiple availability zones and regions, but customers are responsible for designing applications, backups and recovery processes. Without DR planning, a regional outage, ransomware or accidental deletion can still cause major data loss and downtime.

What is the cheapest cloud DR strategy?

Backup and restore is the lowest-cost strategy because little infrastructure runs in the DR location until needed. It has the longest recovery time and largest potential data loss. It suits less critical systems, while revenue-critical systems usually need pilot light, warm standby or active-active.

How often should we test disaster recovery?

Test DR at least once a year for every application tier, and more frequently, such as quarterly, for critical systems. Combine tabletop exercises, component restore tests and periodic full failover tests, and measure results against your RTO and RPO targets.

What is the difference between backup and disaster recovery?

Backup is making copies of data so it can be restored. Disaster recovery is the broader capability to restore entire services, including infrastructure, applications, data, networking and access, within defined time and data-loss targets. Backups are one component of a DR strategy.

Nidhish Joy

Written by

Nidhish Joy

Co-founder & CEO · 10 articles

Nidhish co-founded Crozaint in 2018 and leads it as CEO — 250+ cloud engagements, a 35-strong team running 24/7 operations for 100+ critical applications, and AWS Advanced and Microsoft Gold partner status along the way. He works where technology, strategy and investment meet: AI-first businesses, FinOps and technology economics, and the partnerships that make them real. He is also the first call on any new engagement.

JosephReviewed for technical accuracy by Joseph, Cloud Consulting.

After the reading

Reading About Managed Services Is the Easy Part.Doing It in Your Estate Is Ours.

Thirty minutes with the people who wrote this. We look at your setup, say what we would fix first and leave you with a plan, whether or not you go further with us.

  • A look at your estate, not a demo
  • What we would fix first, and why
  • A plan you keep, whether or not you hire us
Nidhish Joy

Talk to Nidhish

Wrote this article · Co-founder & CEO

Thirty minutes on your estate. Nidhish looks at what you have and tells you what we would do first.

Book 30 Minutes

No deck, no pitch, no commitment.