What are the disaster recovery strategies in AWS?
Compare the four AWS disaster recovery strategies — Backup and Restore, Pilot Light, Warm Standby, and Active/Active — and how RTO and RPO drive your choice.
Expected Interview Answer
AWS defines four disaster recovery strategies along a cost-versus-recovery-speed spectrum: Backup and Restore, Pilot Light, Warm Standby, and Multi-Site Active/Active — each trading higher cost for lower RTO (recovery time) and RPO (data loss).
Backup and Restore is cheapest: you back up data (snapshots, S3, AWS Backup) and rebuild infrastructure after a disaster, giving RTO/RPO of hours. Pilot Light keeps a minimal core (like a replicated database) always running while the rest is provisioned on demand. Warm Standby runs a scaled-down but fully functional copy that you scale up on failover, cutting RTO to minutes. Multi-Site Active/Active serves traffic from multiple Regions simultaneously with near-zero RTO/RPO but the highest cost and complexity. You choose based on the workload's RTO/RPO requirements and budget.
- Clear trade-off between cost and recovery speed
- RTO and RPO targets drive the right strategy choice
- Resilience against Region or AZ-level failures
- Meets compliance and business-continuity requirements
- Enables tested, repeatable failover instead of improvisation
AI Mentor Explanation
DR strategies are like a team's bench depth against injury. Backup and Restore is having no reserves — you recruit and train replacements after someone is hurt, losing weeks. Pilot Light keeps one fit twelfth man ready. Warm Standby is a full second XI training in the nets, ready in minutes. Active/Active is fielding two complete teams so a loss barely registers — costliest, but play never stops.
Step-by-Step Explanation
Step 1
Define RTO and RPO
Establish the maximum tolerable downtime (RTO) and data loss (RPO) per workload — these targets dictate the strategy.
Step 2
Backup and Restore
Back up data with AWS Backup, EBS/RDS snapshots, and cross-Region S3 replication; rebuild infrastructure via IaC when disaster strikes (RTO/RPO in hours).
Step 3
Pilot Light
Keep critical core components (e.g. a replicated database) always running in the recovery Region; provision the rest on failover.
Step 4
Warm Standby
Run a scaled-down but functional copy of the full stack in another Region and scale it up to production size on failover (RTO in minutes).
Step 5
Multi-Site Active/Active
Serve live traffic from multiple Regions with data replication and Route 53 routing for near-zero RTO/RPO at the highest cost.
Step 6
Test and automate failover
Regularly run DR drills, automate with IaC and Route 53 health-check failover, and validate that recovery meets targets.
What Interviewer Expects
- Names all four strategies in order of cost and recovery speed
- Explains RTO and RPO and how they drive the choice
- Distinguishes Pilot Light from Warm Standby precisely
- Mentions cross-Region replication and Route 53 failover
- Stresses testing DR plans rather than assuming they work
Common Mistakes
- Confusing Pilot Light (core always on) with Warm Standby (full scaled-down stack running)
- Never testing the failover, so recovery fails during a real disaster
- Ignoring RPO and losing more data than the business can tolerate
- Choosing Active/Active for workloads that don't justify its cost and complexity
- Backing up data but not automating infrastructure recreation with IaC
Best Answer (HR Friendly)
“AWS offers four disaster recovery options ranging from cheap-but-slow to expensive-but-instant: simple backups you restore later, a small always-on core, a scaled-down running copy, and a full duplicate serving traffic at the same time. You pick based on how quickly the business must recover and how much data loss it can accept.”
Code Example
Resources:
PrimaryRecord:
Type: AWS::Route53::RecordSet
Properties:
HostedZoneId: Z123456ABCDEFG
Name: app.example.com
Type: A
SetIdentifier: primary-us-east-1
Failover: PRIMARY
AliasTarget:
DNSName: primary-alb.us-east-1.elb.amazonaws.com
HostedZoneId: Z35SXDOTRQ7X7K
EvaluateTargetHealth: true
HealthCheckId: !Ref PrimaryHealthCheck
StandbyRecord:
Type: AWS::Route53::RecordSet
Properties:
HostedZoneId: Z123456ABCDEFG
Name: app.example.com
Type: A
SetIdentifier: standby-us-west-2
Failover: SECONDARY
AliasTarget:
DNSName: standby-alb.us-west-2.elb.amazonaws.com
HostedZoneId: Z1H1FL5HABSF5
EvaluateTargetHealth: true
PrimaryHealthCheck:
Type: AWS::Route53::HealthCheck
Properties:
HealthCheckConfig:
Type: HTTPS
FullyQualifiedDomainName: primary-alb.us-east-1.elb.amazonaws.com
ResourcePath: /health
RequestInterval: 30
FailureThreshold: 3Follow-up Questions
- How do RTO and RPO influence which DR strategy you pick?
- What is the exact difference between Pilot Light and Warm Standby?
- How does Route 53 health-check failover automate recovery?
- How would you replicate an RDS database across Regions for DR?
- How do you regularly test a disaster recovery plan without impacting production?
MCQ Practice
1. Which DR strategy has the lowest cost but the highest RTO?
Backup and Restore only stores backups and rebuilds later, making it cheapest but slowest to recover.
2. What does RPO measure?
RPO (Recovery Point Objective) is the maximum amount of data, measured in time, you can afford to lose.
3. Which strategy keeps only critical core components (like a replicated database) always running?
Pilot Light keeps a minimal always-on core and provisions the remaining infrastructure on failover.
Flash Cards
Four AWS DR strategies (cheapest to costliest)? — Backup and Restore, Pilot Light, Warm Standby, Multi-Site Active/Active.
RTO vs RPO? — RTO is how long recovery may take (downtime); RPO is how much data loss (time) is acceptable.
Pilot Light vs Warm Standby? — Pilot Light keeps only a minimal core running; Warm Standby runs a scaled-down full stack ready to scale up.
Which strategy gives near-zero RTO/RPO? — Multi-Site Active/Active, serving traffic from multiple Regions at once — highest cost and complexity.