What you will learn
- What constitutes a disaster in AWS terms — it is not just earthquakes and fires
- RPO and RTO — the two numbers that define your entire DR strategy and budget
- The four DR strategies — Backup and Restore, Pilot Light, Warm Standby, Multi-Site
- Exactly when each strategy makes sense and what it costs relative to alternatives
- AWS Elastic Disaster Recovery (DRS) — continuous block-level replication and the Staging vs Production area design
- AWS Database Migration Service (DMS) — homogeneous and heterogeneous migrations with CDC
- Aurora migration paths — Snapshot, Read Replica, Percona XtraBackup, and mysqldump
- On-premises strategy tools — VM Import, Application Discovery Service, MGN
- AWS Backup — centralised backup across all services, Vault Lock, and WORM protection
- How to choose between internet, Direct Connect, and Snowball for large data transfers
Why this matters
When a production database crashes at 2 AM or a ransomware attack encrypts an entire on-premises data center, the difference between 5 minutes of downtime and 5 days of downtime is whether you designed your DR architecture before the disaster — not during it. At Razorpay, payment processing must resume within minutes of any incident — that SLA requires a Warm Standby or better. At a company with RBI compliance requirements, data must be recoverable within a defined window. These are not theoretical exercises. RPO and RTO are commitments made to the business, and every architecture decision is a direct trade-off between cost and how well those commitments can be kept.
What is a Disaster
Any event that negatively impacts business continuity or finances qualifies as a disaster. It does not have to be an earthquake or a fire.
Developer accidentally deletes a production RDS table → disasterMisconfigured security group exposes database to internet → disasterRegion-level AWS outage → disasterRansomware encrypts on-premises systems → disasterCorrupted deployment brings down the application → disasterRPO and RTO — The Two Numbers That Define Everything
Before picking any DR strategy, two numbers must be confirmed in writing from the business.
RPO (Recovery Point Objective) — how much data loss is acceptable:
Example:2:00 PM → last snapshot taken7:00 PM → disaster hits7:05 PM → you begin restoring from the 2 PM snapshot 5 hours of data is gone forever.Your RPO = 5 hours.RTO (Recovery Time Objective) — how much downtime is acceptable:
Example:7:00 PM → disaster hits, system goes down10:00 PM → system restored and live again 3 hours of users getting errors.Your RTO = 3 hours. Lower RPO = less data loss = more money and infrastructureLower RTO = less downtime = more money and infrastructureBoth require real investment — you are always trading cost against recovery speedDR works across three environments:
On-premise to On-premise → traditional DR, expensive, your own hardwareOn-premise to AWS Cloud → hybrid, on-premise primary, AWS recovery siteAWS Region A to Region B → fully cloud, cross-region active or passiveDR Strategy 1 — Backup and Restore
The cheapest and simplest strategy. Highest RPO and RTO of all four options. Nothing runs in AWS during normal times except stored backups.
Normal time: On-premises servers → Storage Gateway/Snowball → S3 → Glacier (lifecycle) EC2, RDS, Redshift → scheduled snapshots → S3 On disaster: Restore EC2 instances from AMIs stored in S3 Restore databases from RDS snapshots Everything rebuilt from scratch RPO: Hours to days (depends on snapshot frequency)RTO: Hours (full rebuild from scratch)Cost: $ (cheapest)Common MistakeTeams set snapshot frequency to once per day because it is the default, then discover their RPO is actually 24 hours — meaning they can lose an entire day of transactions. Set backup frequency based on your actual RPO requirement, not the default.
DR Strategy 2 — Pilot Light
Faster than Backup and Restore because the most critical and hardest-to-restore component — the database — is always running in AWS and always in sync. Only the application server is stopped.
Why this split? Databases: slow to restore — hold all your data, must be fully recovered first App servers: fast to start — stateless, just need a database to connect to Normal time: On-premises: App Server (running) + Primary DB (running, all traffic) AWS: RDS Secondary (running, silently syncing) + EC2 (STOPPED, ready to launch) On disaster (on-premises goes down): RDS Secondary → already has all data, promoted to Primary EC2 → started (takes minutes, not hours) Route 53 → updated to point to AWS EC2 EC2 connects to RDS → app is live RPO: Minutes to hoursRTO: Minutes to hoursCost: $$ (DB running, EC2 stopped)RememberIn Pilot Light, the DB is always running (most critical, hardest to restore). The EC2 is stopped (cheapest to keep off). Backup and Restore restores everything from scratch. Pilot Light already has the DB live — so recovery is measured in minutes instead of hours.
DR Strategy 3 — Warm Standby
The entire stack runs in AWS at minimum size, not just the database. Nothing needs to start — only scales up. This dramatically reduces RTO.
Normal time: On-premises: App Server (full production) + Primary DB Route 53 → pointing to on-premises AWS: ELB + EC2 Auto Scaling (minimum size, e.g. 1 small instance) RDS Secondary (running, continuously syncing) Route 53 → pointing to on-premises (AWS stack idle) On disaster: Route 53 fails over to AWS EC2 Auto Scaling scales UP to production instance count and size ELB distributes load across scaled EC2 instances RDS Secondary already has all data → takes over as Primary RPO: MinutesRTO: Minutes (scale up, not start up)Cost: $$$ (full stack at minimum size)Pilot Light vs Warm Standby:
Pilot Light: DB running, EC2 stopped. EC2 must START (minutes).Warm Standby: DB running, EC2 running at minimum size. EC2 just SCALES UP (faster).DR Strategy 4 — Multi-Site and Hot Site
The most expensive and fastest strategy. Both on-premises and AWS run at full production scale simultaneously. Route 53 actively sends real traffic to both at the same time.
Normal time (both active): On-premises: App Server (production) + Primary DB AWS: EC2 Auto Scaling (production scale) + RDS Secondary (running) Route 53: active-active → both environments serve real traffic simultaneously On disaster: Route 53 detects on-premises failure All traffic shifts to AWS instantly AWS already at production scale → zero scale-up needed RPO: SecondsRTO: Seconds to minutesCost: $$$$ (two full production environments)All AWS Multi-Region follows the same pattern without on-premises:
AWS Region 1 (primary): AWS Region 2 (secondary): EC2 Auto Scaling (production) EC2 Auto Scaling (production) Aurora Global (primary) Aurora Global (secondary replica) Route 53 (active-active) connecting bothAurora Global handles cross-region replication automatically with sub-second lag.
DR Strategy Comparison:
| Strategy | RPO | RTO | Cost | What runs in AWS normally |
|---|---|---|---|---|
| Backup and Restore | Hours to Days | Hours | $ | Nothing — stored backups only |
| Pilot Light | Minutes to Hours | Minutes to Hours | $$ | DB only (RDS running, EC2 stopped) |
| Warm Standby | Minutes | Minutes | $$$ | Full stack at minimum size |
| Multi-Site | Seconds | Seconds to Minutes | $$$$ | Full stack at production scale |
AWS Elastic Disaster Recovery (DRS)
Previously called CloudEndure Disaster Recovery. Continuously replicates your physical, virtual, or cloud-based servers into AWS so you can recover in minutes when disaster hits.
Why two areas — Staging and Production:
You do not want to pay for full production EC2 instances 24 hours a day just waiting for a disaster. DRS splits into two areas:
Staging area (always running, low cost): Small EC2 + EBS instances Only job: receive and store continuous replication from on-premises Not serving any users, not running your app Like a spare tyre in your boot — always there, updated, costs little Production area (only exists when disaster hits): Full-sized Target EC2 + EBS spun up from Staging data Your actual app runs here and serves users Created on failover, shut down after failback Like the spare tyre going on the wheel — now doing the real jobFull DRS flow:
Normal time: Install AWS Replication Agent on on-premises servers Agent replicates continuously (block-level, seconds lag) → AWS Staging Staging holds a live, up-to-date copy at minimal cost Disaster hits (failover): Trigger failover → DRS spins up full Production EC2 + EBS from Staging data Minutes to recovery — data was already there, nothing to restore Route 53 points to AWS Production → users hit AWS, notice nothing On-premises fixed (failback): DRS replicates Production data back to on-premises in background Once on-premises ready → trigger failback Traffic switches back to on-premises AWS Production shuts down → back to Staging onlyDRS vs MGN:
| DRS | MGN | |
|---|---|---|
| Purpose | Unexpected disaster recovery | Planned permanent migration |
| After complete | Failback to on-premises | On-premises shut down forever |
| On-premises after | Returns to normal | Decommissioned |
AWS Database Migration Service (DMS)
DMS migrates databases to AWS quickly and securely — and the source database stays available and running throughout. No downtime for users. Self-healing — if the migration fails midway, it recovers automatically.
You must create an EC2 instance to run DMS — it sits between source and target and does all the migration work. DMS is not serverless.
Two types of migration:
Homogeneous migration (same DB engine): Oracle → Oracle PostgreSQL → RDS PostgreSQL MySQL → RDS MySQL No schema conversion needed — just migrate the data Heterogeneous migration (different DB engines): MS SQL Server → Aurora Oracle → MySQL Schema conversion needed first → use AWS Schema Conversion Tool (SCT) Then DMS migrates the dataAWS Schema Conversion Tool (SCT):
OLTP: SQL Server or Oracle → MySQL, PostgreSQL, AuroraOLAP: Teradata or Oracle → Amazon RedshiftRememberYou do NOT need SCT if migrating the same engine. On-premises PostgreSQL to RDS PostgreSQL — schema stays identical. SCT is only needed when the engine type itself changes.
CDC — Continuous Data Capture:
DMS uses CDC to continuously capture and replicate every new change from source to target while migration is in progress.
Oracle DB (source) → DMS Replication Instance (EC2) → MySQL DB (target) During migration: Initial bulk load: DMS copies all existing data Ongoing: CDC captures every INSERT, UPDATE, DELETE on source applies them to target in real time Result: target stays in sync with source throughout migrationDMS Multi-AZ Deployment:
AZ-A: DMS Replication Instance → actively migrating dataAZ-B: DMS Standby Replica → synchronously replicating, waiting If AZ-A fails: AZ-B standby takes over automatically, migration continuesAurora Migration Paths
When your source is RDS MySQL or PostgreSQL, two paths exist that do not require DMS.
Option 1 — Snapshot (simplest, one-time move):
Take DB snapshot from RDS MySQL/PostgreSQLRestore directly as Aurora MySQL/PostgreSQLBest when: comfortable with brief cutover window and one-time copyOption 2 — Read Replica (stay in sync during migration):
Create Aurora Read Replica from your RDS instanceWait until replication lag = 0 (Aurora fully caught up)Promote Aurora as independent DB clusterBest when: need data in sync throughout, cannot afford any data lossExternal MySQL to Aurora:
Option A — Percona XtraBackup via S3 (faster, large DBs): Run Percona XtraBackup on your server Upload backup file to S3 Import directly into Aurora MySQL Use for: large production databases, physical backup is 3-10x faster than logical Option B — mysqldump (simpler, small DBs): Export data directly into Aurora MySQL Logical export — slower for large databases Use for: databases under 10 GB, schema-only backups| Scenario | Best option |
|---|---|
| Both source and target live simultaneously | DMS with CDC |
| RDS MySQL → Aurora, fast one-time move | Snapshot restore |
| RDS MySQL → Aurora, stay in sync during migration | Read Replica → promote |
| External MySQL → Aurora, large database | Percona XtraBackup via S3 |
| External MySQL → Aurora, small database | mysqldump |
AWS Backup — Centralised Backup Management
Before AWS Backup you configured backups inside each service separately — RDS had its own backup settings, EBS had snapshots, DynamoDB had its own configuration. No central visibility or policy.
AWS Backup centralises all of this — one managed service to automate and manage backups across every supported service.
Supported services:
EC2 / EBS, S3, RDS (all engines) / Aurora / DynamoDBDocumentDB, Neptune, EFS, FSx, Storage Gateway (Volume Gateway)Backup Plans — what you configure:
| Setting | Options |
|---|---|
| Backup frequency | Every 12 hours, daily, weekly, monthly, or cron |
| Backup window | Time of day to run |
| Transition to Cold Storage | Never, Days, Weeks, Months, Years |
| Retention period | Days, Weeks, Months, Years, or Always |
Tag-based backup policies:
Assign a tag (e.g. backup: daily) to any resource and AWS Backup automatically includes it in the matching plan. No need to manually add each resource.
Additional features:
PITR — restore to any specific moment for supported servicesCross-region backups — store backups in a different region for DRCross-account backups — store in a completely separate AWS accountAWS Backup Vault Lock:
A backup that can be deleted is not a real backup. Ransomware attackers target backup systems first precisely because deleting backups forces companies to pay.
WORM (Write Once Read Many) state on all backupsOnce written, cannot be deleted or modified by anyoneEven the root user cannot delete backups when Vault Lock is enabledSecurityBackup Vault Lock is one of the very few AWS controls that even the root user cannot override. Once enabled, plan your retention policies carefully before turning it on — you cannot shorten or change them after.
Transferring Large Amounts of Data into AWS
Real scenario: transfer 200 TB from on-premises into AWS. You have a 100 Mbps internet connection.
Option 1 — Internet or Site-to-Site VPN:
200 TB at 100 Mbps:200 × 1000 (GB) × 1000 (MB) × 8 (Mb) / 100 Mbps = 16,000,000 seconds = 185 days Result: 185 days → not viable for large one-time transfersOption 2 — Direct Connect (1 Gbps):
200 TB at 1 Gbps:200 × 1000 × 1000 × 8 / 1000 Mbps = 1,600,000 seconds = 18.5 daysBut Direct Connect takes over a month just to set up the physical connection Result: faster than internet, but setup delay makes it wrong for urgent one-time movesOption 3 — Snowball:
AWS ships a physical deviceLoad data locally at full disk speed — no network bottleneckShip device back to AWS → data ingested into S3End-to-end: about 1 week Result: 1 week total → correct choice for large one-time data transferFor ongoing replication after the initial transfer, use Site-to-Site VPN or Direct Connect with DMS or DataSync.
Quick decision:
| Scenario | Best option |
|---|---|
| One-time large transfer (TBs) | Snowball |
| Ongoing replication or continuous sync | Site-to-Site VPN or Direct Connect + DMS |
| Small data, need it quickly | Internet / VPN |
| Large data, can wait for connection setup | Direct Connect |
Hands-on Lab — AWS Backup with Vault Lock and Cross-Region Snapshot
Step 1 — Create a Backup Vault
AWS Backup → Backup vaults → Create backup vaultVault name: devops-prod-vaultEncryption key: AWS managed key (or your CMK for compliance)Create backup vaultStep 2 — Create a Backup Plan
AWS Backup → Backup plans → Create backup planStart with a template: Daily-35day-RetentionPlan name: devops-daily-backupCreate plan This creates:Daily backup at 5 AM UTCRetain backups for 35 daysStore in devops-prod-vaultStep 3 — Assign resources to the backup plan
devops-daily-backup → Assign resourcesResource assignment name: rds-productionIAM role: Default role (AWSBackupDefaultServiceRole)Resource selection: Include specific resource typesResource type: RDS → select your RDS instanceAssign resourcesStep 4 — Run an on-demand backup
AWS Backup → Protected resources → select your RDS instanceCreate on-demand backupBackup vault: devops-prod-vaultRetention: 7 daysCreate on-demand backup Backup Jobs → watch status change from Created → Running → CompletedStep 5 — Enable Vault Lock (WORM protection)
AWS Backup → Backup vaults → devops-prod-vaultVault lock → EditMinimum retention days: 7Maximum retention days: 365Enable vault lock → confirm Once locked, no backup in this vault can be deleted before 7 days.Even root cannot override this — this is the WORM guarantee.Step 6 — Enable CloudTrail to audit all backup actions
AWS Backup → all create, delete, and restore actions appear in CloudTrailCloudTrail → Event history → filter EventSource: backup.amazonaws.comEvery backup operation is now auditable — who triggered it, when, result.Step 7 — Cleanup
AWS Backup → Backup jobs → wait for jobs to completeProtected resources → deselect your RDS instanceBackup vaults → devops-prod-vault → Recovery points → delete all pointsBackup vaults → devops-prod-vault → Delete vaultBackup plans → devops-daily-backup → Delete planProduction Best Practices and Common Pitfalls
- Choose the DR strategy based on confirmed RPO and RTO numbers — never design DR without these numbers agreed in writing with the business
- Test your DR plan quarterly — run the full failover procedure including EC2 startup, database connectivity, and application health checks
- Use AWS Backup instead of per-service manual backup configuration — central visibility prevents the common mistake of some services having no backup policy
- Enable Backup Vault Lock for compliance-sensitive workloads — ransomware that cannot delete your backups cannot force you to pay
- Use Snowball for large initial data transfers to AWS — attempting 100 TB over the internet wastes months
- Run the Pilot Light EC2 in a stopped state but test startup time regularly — a stopped EC2 with a complex startup script might take 20 minutes instead of 2
Quick Reference and Troubleshooting Commands
| Task | Command |
|---|---|
| List backup plans | aws backup list-backup-plans --region ap-south-1 |
| Start on-demand backup | aws backup start-backup-job --backup-vault-name Default --resource-arn <arn> --iam-role-arn <role> |
| List backup jobs | aws backup list-backup-jobs --region ap-south-1 |
| List DMS replication instances | aws dms describe-replication-instances --region ap-south-1 |
| Describe RDS replication lag | aws rds describe-db-instances --db-instance-identifier <id> --query 'DBInstances[0].StatusInfos' |
| List Snowball jobs | aws snowball list-jobs |
| Check MGN replication status | aws mgn describe-source-servers --region ap-south-1 |
Common problems and fixes:
| Problem | Likely cause | Fix |
|---|---|---|
| RPO exceeding SLA | Snapshot frequency too low | Increase backup frequency or enable continuous replication via DRS |
| RTO exceeding SLA | Wrong DR strategy chosen | Upgrade from Backup and Restore to Pilot Light or Warm Standby |
| DMS migration losing data | CDC not enabled from start of full load | Enable CDC before full load begins — overlap ensures no write is missed |
| Pilot Light EC2 taking too long to start | Complex startup script | Pre-bake AMI with application installed, test startup time quarterly |
| Vault Lock preventing needed backup deletion | Retention period too long | Test retention values thoroughly before applying Vault Lock — it is irreversible |
Common MistakeChoosing Backup and Restore when the actual RPO requirement is 30 minutes. Daily snapshots give RPO of up to 24 hours. If the business needs 30-minute RPO, you need Pilot Light with a continuously syncing database — not daily snapshots in S3. Always get RPO and RTO confirmed in writing before designing anything.
Common MistakeEnabling Backup Vault Lock without testing retention policies first. Once WORM is applied, nobody — not even root — can delete a backup before the minimum retention period. If you set a 7-day minimum and need to delete a backup for compliance reasons before then, you cannot. Test in a non-production vault first.
SecurityDMS requires an EC2 instance to run and is billed for the entire time migration is running. For long-running CDC migrations lasting weeks or months, this EC2 cost adds up. Size the Replication Instance correctly for your data volume — too small causes migration lag and data accumulation, too large wastes money.
Common Mistakes to Avoid
Common MistakeSetting automated backup retention to 0 days. This disables PITR entirely. If a developer accidentally truncates a production table, there is nothing to restore from. Always keep retention at minimum 7 days for production databases.
Common MistakeChoosing Backup and Restore when RPO is 30 minutes. Daily snapshots give up to 24 hours of data loss. For 30-minute RPO you need Pilot Light with a continuously syncing database. Always confirm RPO and RTO numbers in writing before designing anything.
SecurityEnable Backup Vault Lock for compliance workloads. Ransomware attackers target backups first — deleting your backups forces you to pay. Vault Lock WORM protection means even root cannot delete backups before the minimum retention period.