Skip to main content

AWS Disaster Recovery - Backup Restore, Pilot Light, Warm Standby, Multi-Site

Design and implement the four AWS disaster recovery strategies based on RPO and RTO requirements, from simple S3 backups to active-active multi-site.

What you will learn

  • What constitutes a disaster in AWS terms — it is not just earthquakes and fires
  • RPO and RTO — the two numbers that define your entire DR strategy and budget
  • The four DR strategies — Backup and Restore, Pilot Light, Warm Standby, Multi-Site
  • Exactly when each strategy makes sense and what it costs relative to alternatives
  • AWS Elastic Disaster Recovery (DRS) — continuous block-level replication and the Staging vs Production area design
  • AWS Database Migration Service (DMS) — homogeneous and heterogeneous migrations with CDC
  • Aurora migration paths — Snapshot, Read Replica, Percona XtraBackup, and mysqldump
  • On-premises strategy tools — VM Import, Application Discovery Service, MGN
  • AWS Backup — centralised backup across all services, Vault Lock, and WORM protection
  • How to choose between internet, Direct Connect, and Snowball for large data transfers

Why this matters

When a production database crashes at 2 AM or a ransomware attack encrypts an entire on-premises data center, the difference between 5 minutes of downtime and 5 days of downtime is whether you designed your DR architecture before the disaster — not during it. At Razorpay, payment processing must resume within minutes of any incident — that SLA requires a Warm Standby or better. At a company with RBI compliance requirements, data must be recoverable within a defined window. These are not theoretical exercises. RPO and RTO are commitments made to the business, and every architecture decision is a direct trade-off between cost and how well those commitments can be kept.

What is a Disaster

Any event that negatively impacts business continuity or finances qualifies as a disaster. It does not have to be an earthquake or a fire.

◈ DIAGRAM
Developer accidentally deletes a production RDS table → disaster
Misconfigured security group exposes database to internet → disaster
Region-level AWS outage → disaster
Ransomware encrypts on-premises systems → disaster
Corrupted deployment brings down the application → disaster

RPO and RTO — The Two Numbers That Define Everything

Before picking any DR strategy, two numbers must be confirmed in writing from the business.

RPO (Recovery Point Objective) — how much data loss is acceptable:

◈ DIAGRAM
Example:
2:00 PM → last snapshot taken
7:00 PM → disaster hits
7:05 PM → you begin restoring from the 2 PM snapshot
5 hours of data is gone forever.
Your RPO = 5 hours.

RTO (Recovery Time Objective) — how much downtime is acceptable:

◈ DIAGRAM
Example:
7:00 PM → disaster hits, system goes down
10:00 PM → system restored and live again
3 hours of users getting errors.
Your RTO = 3 hours.
Lower RPO = less data loss = more money and infrastructure
Lower RTO = less downtime = more money and infrastructure
Both require real investment — you are always trading cost against recovery speed

DR works across three environments:

◈ DIAGRAM
On-premise to On-premise → traditional DR, expensive, your own hardware
On-premise to AWS Cloud → hybrid, on-premise primary, AWS recovery site
AWS Region A to Region B → fully cloud, cross-region active or passive

DR Strategy 1 — Backup and Restore

The cheapest and simplest strategy. Highest RPO and RTO of all four options. Nothing runs in AWS during normal times except stored backups.

◈ DIAGRAM
Normal time:
On-premises servers → Storage Gateway/Snowball → S3 → Glacier (lifecycle)
EC2, RDS, Redshift → scheduled snapshots → S3
On disaster:
Restore EC2 instances from AMIs stored in S3
Restore databases from RDS snapshots
Everything rebuilt from scratch
RPO: Hours to days (depends on snapshot frequency)
RTO: Hours (full rebuild from scratch)
Cost: $ (cheapest)
Common Mistake

Teams set snapshot frequency to once per day because it is the default, then discover their RPO is actually 24 hours — meaning they can lose an entire day of transactions. Set backup frequency based on your actual RPO requirement, not the default.

DR Strategy 2 — Pilot Light

Faster than Backup and Restore because the most critical and hardest-to-restore component — the database — is always running in AWS and always in sync. Only the application server is stopped.

◈ DIAGRAM
Why this split?
Databases: slow to restore — hold all your data, must be fully recovered first
App servers: fast to start — stateless, just need a database to connect to
Normal time:
On-premises: App Server (running) + Primary DB (running, all traffic)
AWS: RDS Secondary (running, silently syncing) + EC2 (STOPPED, ready to launch)
On disaster (on-premises goes down):
RDS Secondary → already has all data, promoted to Primary
EC2 → started (takes minutes, not hours)
Route 53 → updated to point to AWS EC2
EC2 connects to RDS → app is live
RPO: Minutes to hours
RTO: Minutes to hours
Cost: $$ (DB running, EC2 stopped)
Remember

In Pilot Light, the DB is always running (most critical, hardest to restore). The EC2 is stopped (cheapest to keep off). Backup and Restore restores everything from scratch. Pilot Light already has the DB live — so recovery is measured in minutes instead of hours.

DR Strategy 3 — Warm Standby

The entire stack runs in AWS at minimum size, not just the database. Nothing needs to start — only scales up. This dramatically reduces RTO.

◈ DIAGRAM
Normal time:
On-premises: App Server (full production) + Primary DB
Route 53 → pointing to on-premises
AWS: ELB + EC2 Auto Scaling (minimum size, e.g. 1 small instance)
RDS Secondary (running, continuously syncing)
Route 53 → pointing to on-premises (AWS stack idle)
On disaster:
Route 53 fails over to AWS
EC2 Auto Scaling scales UP to production instance count and size
ELB distributes load across scaled EC2 instances
RDS Secondary already has all data → takes over as Primary
RPO: Minutes
RTO: Minutes (scale up, not start up)
Cost: $$$ (full stack at minimum size)

Pilot Light vs Warm Standby:

TEXT
Pilot Light: DB running, EC2 stopped. EC2 must START (minutes).
Warm Standby: DB running, EC2 running at minimum size. EC2 just SCALES UP (faster).

DR Strategy 4 — Multi-Site and Hot Site

The most expensive and fastest strategy. Both on-premises and AWS run at full production scale simultaneously. Route 53 actively sends real traffic to both at the same time.

◈ DIAGRAM
Normal time (both active):
On-premises: App Server (production) + Primary DB
AWS: EC2 Auto Scaling (production scale) + RDS Secondary (running)
Route 53: active-active → both environments serve real traffic simultaneously
On disaster:
Route 53 detects on-premises failure
All traffic shifts to AWS instantly
AWS already at production scale → zero scale-up needed
RPO: Seconds
RTO: Seconds to minutes
Cost: $$$$ (two full production environments)

All AWS Multi-Region follows the same pattern without on-premises:

TEXT
AWS Region 1 (primary): AWS Region 2 (secondary):
EC2 Auto Scaling (production) EC2 Auto Scaling (production)
Aurora Global (primary) Aurora Global (secondary replica)
Route 53 (active-active) connecting both

Aurora Global handles cross-region replication automatically with sub-second lag.

DR Strategy Comparison:

Strategy RPO RTO Cost What runs in AWS normally
Backup and Restore Hours to Days Hours $ Nothing — stored backups only
Pilot Light Minutes to Hours Minutes to Hours $$ DB only (RDS running, EC2 stopped)
Warm Standby Minutes Minutes $$$ Full stack at minimum size
Multi-Site Seconds Seconds to Minutes $$$$ Full stack at production scale

AWS Elastic Disaster Recovery (DRS)

Previously called CloudEndure Disaster Recovery. Continuously replicates your physical, virtual, or cloud-based servers into AWS so you can recover in minutes when disaster hits.

Why two areas — Staging and Production:

You do not want to pay for full production EC2 instances 24 hours a day just waiting for a disaster. DRS splits into two areas:

TEXT
Staging area (always running, low cost):
Small EC2 + EBS instances
Only job: receive and store continuous replication from on-premises
Not serving any users, not running your app
Like a spare tyre in your boot — always there, updated, costs little
Production area (only exists when disaster hits):
Full-sized Target EC2 + EBS spun up from Staging data
Your actual app runs here and serves users
Created on failover, shut down after failback
Like the spare tyre going on the wheel — now doing the real job

Full DRS flow:

◈ DIAGRAM
Normal time:
Install AWS Replication Agent on on-premises servers
Agent replicates continuously (block-level, seconds lag) → AWS Staging
Staging holds a live, up-to-date copy at minimal cost
Disaster hits (failover):
Trigger failover → DRS spins up full Production EC2 + EBS from Staging data
Minutes to recovery — data was already there, nothing to restore
Route 53 points to AWS Production → users hit AWS, notice nothing
On-premises fixed (failback):
DRS replicates Production data back to on-premises in background
Once on-premises ready → trigger failback
Traffic switches back to on-premises
AWS Production shuts down → back to Staging only

DRS vs MGN:

DRS MGN
Purpose Unexpected disaster recovery Planned permanent migration
After complete Failback to on-premises On-premises shut down forever
On-premises after Returns to normal Decommissioned

AWS Database Migration Service (DMS)

DMS migrates databases to AWS quickly and securely — and the source database stays available and running throughout. No downtime for users. Self-healing — if the migration fails midway, it recovers automatically.

You must create an EC2 instance to run DMS — it sits between source and target and does all the migration work. DMS is not serverless.

Two types of migration:

◈ DIAGRAM
Homogeneous migration (same DB engine):
Oracle → Oracle
PostgreSQL → RDS PostgreSQL
MySQL → RDS MySQL
No schema conversion needed — just migrate the data
Heterogeneous migration (different DB engines):
MS SQL Server → Aurora
Oracle → MySQL
Schema conversion needed first → use AWS Schema Conversion Tool (SCT)
Then DMS migrates the data

AWS Schema Conversion Tool (SCT):

◈ DIAGRAM
OLTP: SQL Server or Oracle → MySQL, PostgreSQL, Aurora
OLAP: Teradata or Oracle → Amazon Redshift
Remember

You do NOT need SCT if migrating the same engine. On-premises PostgreSQL to RDS PostgreSQL — schema stays identical. SCT is only needed when the engine type itself changes.

CDC — Continuous Data Capture:

DMS uses CDC to continuously capture and replicate every new change from source to target while migration is in progress.

◈ DIAGRAM
Oracle DB (source) → DMS Replication Instance (EC2) → MySQL DB (target)
During migration:
Initial bulk load: DMS copies all existing data
Ongoing: CDC captures every INSERT, UPDATE, DELETE on source
applies them to target in real time
Result: target stays in sync with source throughout migration

DMS Multi-AZ Deployment:

◈ DIAGRAM
AZ-A: DMS Replication Instance → actively migrating data
AZ-B: DMS Standby Replica → synchronously replicating, waiting
If AZ-A fails: AZ-B standby takes over automatically, migration continues

Aurora Migration Paths

When your source is RDS MySQL or PostgreSQL, two paths exist that do not require DMS.

Option 1 — Snapshot (simplest, one-time move):

TEXT
Take DB snapshot from RDS MySQL/PostgreSQL
Restore directly as Aurora MySQL/PostgreSQL
Best when: comfortable with brief cutover window and one-time copy

Option 2 — Read Replica (stay in sync during migration):

TEXT
Create Aurora Read Replica from your RDS instance
Wait until replication lag = 0 (Aurora fully caught up)
Promote Aurora as independent DB cluster
Best when: need data in sync throughout, cannot afford any data loss

External MySQL to Aurora:

TEXT
Option A — Percona XtraBackup via S3 (faster, large DBs):
Run Percona XtraBackup on your server
Upload backup file to S3
Import directly into Aurora MySQL
Use for: large production databases, physical backup is 3-10x faster than logical
Option B — mysqldump (simpler, small DBs):
Export data directly into Aurora MySQL
Logical export — slower for large databases
Use for: databases under 10 GB, schema-only backups
Scenario Best option
Both source and target live simultaneously DMS with CDC
RDS MySQL → Aurora, fast one-time move Snapshot restore
RDS MySQL → Aurora, stay in sync during migration Read Replica → promote
External MySQL → Aurora, large database Percona XtraBackup via S3
External MySQL → Aurora, small database mysqldump

AWS Backup — Centralised Backup Management

Before AWS Backup you configured backups inside each service separately — RDS had its own backup settings, EBS had snapshots, DynamoDB had its own configuration. No central visibility or policy.

AWS Backup centralises all of this — one managed service to automate and manage backups across every supported service.

Supported services:

TEXT
EC2 / EBS, S3, RDS (all engines) / Aurora / DynamoDB
DocumentDB, Neptune, EFS, FSx, Storage Gateway (Volume Gateway)

Backup Plans — what you configure:

Setting Options
Backup frequency Every 12 hours, daily, weekly, monthly, or cron
Backup window Time of day to run
Transition to Cold Storage Never, Days, Weeks, Months, Years
Retention period Days, Weeks, Months, Years, or Always

Tag-based backup policies:

Assign a tag (e.g. backup: daily) to any resource and AWS Backup automatically includes it in the matching plan. No need to manually add each resource.

Additional features:

TEXT
PITR — restore to any specific moment for supported services
Cross-region backups — store backups in a different region for DR
Cross-account backups — store in a completely separate AWS account

AWS Backup Vault Lock:

A backup that can be deleted is not a real backup. Ransomware attackers target backup systems first precisely because deleting backups forces companies to pay.

TEXT
WORM (Write Once Read Many) state on all backups
Once written, cannot be deleted or modified by anyone
Even the root user cannot delete backups when Vault Lock is enabled
Security

Backup Vault Lock is one of the very few AWS controls that even the root user cannot override. Once enabled, plan your retention policies carefully before turning it on — you cannot shorten or change them after.

Transferring Large Amounts of Data into AWS

Real scenario: transfer 200 TB from on-premises into AWS. You have a 100 Mbps internet connection.

Option 1 — Internet or Site-to-Site VPN:

◈ DIAGRAM
200 TB at 100 Mbps:
200 × 1000 (GB) × 1000 (MB) × 8 (Mb) / 100 Mbps = 16,000,000 seconds = 185 days
Result: 185 days → not viable for large one-time transfers

Option 2 — Direct Connect (1 Gbps):

TEXT
200 TB at 1 Gbps:
200 × 1000 × 1000 × 8 / 1000 Mbps = 1,600,000 seconds = 18.5 days
But Direct Connect takes over a month just to set up the physical connection
Result: faster than internet, but setup delay makes it wrong for urgent one-time moves

Option 3 — Snowball:

◈ DIAGRAM
AWS ships a physical device
Load data locally at full disk speed — no network bottleneck
Ship device back to AWS → data ingested into S3
End-to-end: about 1 week
Result: 1 week total → correct choice for large one-time data transfer

For ongoing replication after the initial transfer, use Site-to-Site VPN or Direct Connect with DMS or DataSync.

Quick decision:

Scenario Best option
One-time large transfer (TBs) Snowball
Ongoing replication or continuous sync Site-to-Site VPN or Direct Connect + DMS
Small data, need it quickly Internet / VPN
Large data, can wait for connection setup Direct Connect

Hands-on Lab — AWS Backup with Vault Lock and Cross-Region Snapshot

Step 1 — Create a Backup Vault

◈ DIAGRAM
AWS Backup → Backup vaults → Create backup vault
Vault name: devops-prod-vault
Encryption key: AWS managed key (or your CMK for compliance)
Create backup vault

Step 2 — Create a Backup Plan

◈ DIAGRAM
AWS Backup → Backup plans → Create backup plan
Start with a template: Daily-35day-Retention
Plan name: devops-daily-backup
Create plan
This creates:
Daily backup at 5 AM UTC
Retain backups for 35 days
Store in devops-prod-vault

Step 3 — Assign resources to the backup plan

◈ DIAGRAM
devops-daily-backup → Assign resources
Resource assignment name: rds-production
IAM role: Default role (AWSBackupDefaultServiceRole)
Resource selection: Include specific resource types
Resource type: RDS → select your RDS instance
Assign resources

Step 4 — Run an on-demand backup

◈ DIAGRAM
AWS Backup → Protected resources → select your RDS instance
Create on-demand backup
Backup vault: devops-prod-vault
Retention: 7 days
Create on-demand backup
Backup Jobs → watch status change from Created → Running → Completed

Step 5 — Enable Vault Lock (WORM protection)

◈ DIAGRAM
AWS Backup → Backup vaults → devops-prod-vault
Vault lock → Edit
Minimum retention days: 7
Maximum retention days: 365
Enable vault lock → confirm
Once locked, no backup in this vault can be deleted before 7 days.
Even root cannot override this — this is the WORM guarantee.

Step 6 — Enable CloudTrail to audit all backup actions

◈ DIAGRAM
AWS Backup → all create, delete, and restore actions appear in CloudTrail
CloudTrail → Event history → filter EventSource: backup.amazonaws.com
Every backup operation is now auditable — who triggered it, when, result.

Step 7 — Cleanup

◈ DIAGRAM
AWS Backup → Backup jobs → wait for jobs to complete
Protected resources → deselect your RDS instance
Backup vaults → devops-prod-vault → Recovery points → delete all points
Backup vaults → devops-prod-vault → Delete vault
Backup plans → devops-daily-backup → Delete plan

Production Best Practices and Common Pitfalls

  • Choose the DR strategy based on confirmed RPO and RTO numbers — never design DR without these numbers agreed in writing with the business
  • Test your DR plan quarterly — run the full failover procedure including EC2 startup, database connectivity, and application health checks
  • Use AWS Backup instead of per-service manual backup configuration — central visibility prevents the common mistake of some services having no backup policy
  • Enable Backup Vault Lock for compliance-sensitive workloads — ransomware that cannot delete your backups cannot force you to pay
  • Use Snowball for large initial data transfers to AWS — attempting 100 TB over the internet wastes months
  • Run the Pilot Light EC2 in a stopped state but test startup time regularly — a stopped EC2 with a complex startup script might take 20 minutes instead of 2

Quick Reference and Troubleshooting Commands

Task Command
List backup plans aws backup list-backup-plans --region ap-south-1
Start on-demand backup aws backup start-backup-job --backup-vault-name Default --resource-arn <arn> --iam-role-arn <role>
List backup jobs aws backup list-backup-jobs --region ap-south-1
List DMS replication instances aws dms describe-replication-instances --region ap-south-1
Describe RDS replication lag aws rds describe-db-instances --db-instance-identifier <id> --query 'DBInstances[0].StatusInfos'
List Snowball jobs aws snowball list-jobs
Check MGN replication status aws mgn describe-source-servers --region ap-south-1

Common problems and fixes:

Problem Likely cause Fix
RPO exceeding SLA Snapshot frequency too low Increase backup frequency or enable continuous replication via DRS
RTO exceeding SLA Wrong DR strategy chosen Upgrade from Backup and Restore to Pilot Light or Warm Standby
DMS migration losing data CDC not enabled from start of full load Enable CDC before full load begins — overlap ensures no write is missed
Pilot Light EC2 taking too long to start Complex startup script Pre-bake AMI with application installed, test startup time quarterly
Vault Lock preventing needed backup deletion Retention period too long Test retention values thoroughly before applying Vault Lock — it is irreversible
Common Mistake

Choosing Backup and Restore when the actual RPO requirement is 30 minutes. Daily snapshots give RPO of up to 24 hours. If the business needs 30-minute RPO, you need Pilot Light with a continuously syncing database — not daily snapshots in S3. Always get RPO and RTO confirmed in writing before designing anything.

Common Mistake

Enabling Backup Vault Lock without testing retention policies first. Once WORM is applied, nobody — not even root — can delete a backup before the minimum retention period. If you set a 7-day minimum and need to delete a backup for compliance reasons before then, you cannot. Test in a non-production vault first.

Security

DMS requires an EC2 instance to run and is billed for the entire time migration is running. For long-running CDC migrations lasting weeks or months, this EC2 cost adds up. Size the Replication Instance correctly for your data volume — too small causes migration lag and data accumulation, too large wastes money.

Common Mistakes to Avoid

Common Mistake

Setting automated backup retention to 0 days. This disables PITR entirely. If a developer accidentally truncates a production table, there is nothing to restore from. Always keep retention at minimum 7 days for production databases.

Common Mistake

Choosing Backup and Restore when RPO is 30 minutes. Daily snapshots give up to 24 hours of data loss. For 30-minute RPO you need Pilot Light with a continuously syncing database. Always confirm RPO and RTO numbers in writing before designing anything.

Security

Enable Backup Vault Lock for compliance workloads. Ransomware attackers target backups first — deleting your backups forces you to pay. Vault Lock WORM protection means even root cannot delete backups before the minimum retention period.

Resources

AWS Direct Connect vs Site-to-Site VPN Failover

AWS Direct Connect vs Site-to-Site VPN Failover

Direct Connect vs VPN isn't really either/or for production — it's a primary-plus-failover pattern. Here's how to design it, and when either/or is right.

5 min read•Aug 2026
Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot, at the crossover where Lambda stops being cheaper — 2026 pricing, invocation thresholds, and interruption math.

5 min read•Aug 2026
Secrets Manager vs Parameter Store vs Vault

Secrets Manager vs Parameter Store vs Vault

AWS Secrets Manager, Parameter Store, and HashiCorp Vault compared for 2026 - cost math, rotation, multi-cloud fit, and the Vault-to-OpenBao fork.

5 min read•Aug 2026
AWS VPC Security: Hardening Every Layer

AWS VPC Security: Hardening Every Layer

Most cloud security incidents start with a misconfigured VPC. Here's how to harden every layer — subnets, Security Groups, NACLs, and IAM — for production.

5 min read•Jul 2026
Event-Driven Architecture on AWS Explained

Event-Driven Architecture on AWS Explained

Event-driven architecture on AWS decouples services and absorbs traffic spikes using SQS, SNS, EventBridge, and Lambda — workflows that scale themselves.

5 min read•Jul 2026
S3 vs RDS vs DynamoDB: Choosing AWS Storage

S3 vs RDS vs DynamoDB: Choosing AWS Storage

Choosing S3, RDS, or DynamoDB wrong costs you in performance, cost, and scalability. Here is a practical decision guide based on your actual access patterns.

5 min read•Jul 2026
AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS bills surprise teams every month. Here are the 8 concrete actions that cut cloud spend by 40-60% without touching your application architecture.

5 min read•Jul 2026
EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2, Lambda, or Fargate — choosing the wrong AWS compute option costs you money and performance. Here is exactly when to use each one in production.

5 min read•Jul 2026

Explore More in AWS Storage and Databases

All 6 Topics

Frequently Asked Questions

Is AWS Disaster Recovery - Backup Restore, Pilot Light, Warm Standby, Multi-Site free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the AWS Disaster Recovery - Backup Restore, Pilot Light, Warm Standby, Multi-Site topic cover?

Design and implement the four AWS disaster recovery strategies based on RPO and RTO requirements, from simple S3 backups to active-active multi-site.