What you will learn
- What RDS is and which database engines it manages for you
- Read Replicas vs Multi-AZ — the most important distinction in RDS and why they solve completely different problems
- Storage Auto Scaling and why running out of disk causes immediate failures
- How to encrypt RDS and authenticate with IAM instead of passwords
- Aurora architecture — what makes it fundamentally different from standard RDS
- Aurora Serverless for workloads that need to scale to zero
- Aurora Global for sub-second cross-region replication
- Aurora Database Cloning — production data in staging in minutes
- RDS backup strategies — automated backups vs manual snapshots
- ElastiCache as a caching layer to take load off RDS
Why this matters
Zerodha processes millions of stock transactions every day. Every trade, every ledger entry, every user position must be written atomically and consistently — this is what relational databases are built for. Razorpay's payment records need relational integrity across users, merchants, and transactions. Swiggy's order service uses Aurora because it needs the performance and reliability of a cloud-native database engine with the familiarity of MySQL.
Choosing the wrong database engine, forgetting to enable Multi-AZ, or not understanding the difference between Read Replicas and Multi-AZ are decisions that teams regret the moment their first production database incident hits.
What is Amazon RDS
RDS (Relational Database Service) is a managed service that runs standard SQL databases on AWS. The key word is managed — AWS handles the parts nobody wants to handle:
AWS manages: Provisioning the underlying EC2 and storage OS patching and maintenance Automated backups Multi-AZ replication and failover Minor version upgrades (optional) Monitoring and CloudWatch metrics You manage: Schema design and queries Application-level optimisation Security Groups and network access Major version upgrades Who has database credentialsYou cannot SSH into the underlying RDS server. That is the trade-off. You give up OS-level access and get a fully managed, highly available database in return.
Supported engines:
PostgreSQL → most popular for new production workloadsMySQL → most widely used overallMariaDB → open-source MySQL-compatibleOracle → enterprise, bring-your-own-licenseMicrosoft SQL → Windows-centric enterpriseAurora → AWS's own engine (covered in detail below)Read Replicas vs Multi-AZ — The Most Important Distinction
This is the concept that every RDS question in production and in exams comes back to. Read Replicas and Multi-AZ solve completely different problems. They are often used together.
Read Replicas — for scaling reads
A Read Replica is a copy of your database that can serve read traffic. It uses asynchronous replication — changes flow from the primary to the replica, but there can be a small lag.
Purpose: Scale out read traffic across multiple copiesReplication: ASYNC — replica is eventually consistent with primaryWho serves: Replica actively serves read queries from your applicationMax replicas: 5 for standard RDS, 15 for AuroraFailover: NOT automatic — you must manually promote a replica to primary Without Read Replica: All reads and writes → single RDS primary Heavy reporting query running at 2 PM → slows down production writes With Read Replica: Writes + critical reads → RDS Primary Reporting queries, analytics → Read Replica (no impact on primary)Cross-region Read Replicas are also possible — useful when you want users in another region to read from a nearby database, or when you need disaster recovery in another region.
Multi-AZ — for disaster recovery
Multi-AZ keeps a standby copy of your database in a different Availability Zone using synchronous replication. The standby serves zero traffic during normal operations — it exists only to take over when the primary fails.
Purpose: High availability and automatic failoverReplication: SYNC — standby is always identical to primary (no lag)Who serves: Standby serves ZERO traffic during normal operationsFailover: Automatic — 60 to 120 seconds, no human action needed Normal state: All traffic (reads AND writes) → RDS Primary in AZ-1a RDS Standby in AZ-1b → receives SYNC replication, does nothing else Primary fails: AWS detects failure DNS endpoint automatically points to the standby Your application reconnects to the new primary Total downtime: 60 to 120 seconds Zero data loss — SYNC replication means standby was always currentRememberRead Replicas = read scaling via ASYNC replication. Multi-AZ = disaster recovery via SYNC replication. A Read Replica can have replication lag. Multi-AZ standby is always identical. Multi-AZ standby cannot serve read traffic. Read Replicas can. They solve different problems. Use both in production.
Can you promote a Read Replica to Multi-AZ?
Not directly. They are separate features. In production your setup typically looks like:
RDS Primary (Multi-AZ enabled) → SYNC to RDS Standby in different AZ (serves no traffic) → ASYNC to Read Replica (serves read traffic, possibly in another region)Storage Auto Scaling
Without Auto Scaling, when your RDS storage fills up, writes start failing immediately. Applications crash. Users see errors. Storage full is not a warning — it is an outage.
With Storage Auto Scaling, RDS detects when free space is running low and automatically expands storage — no downtime, no manual action.
You set: Maximum Storage Threshold (e.g. 100 GB)RDS monitors: free storage continuouslyTrigger: free space < 10% AND low for 5 minutes AND 6 hours since last expansionAction: storage expands automatically in incrementsTipEnable Storage Auto Scaling on every production RDS instance from day one. Storage running out is an avoidable failure mode. The auto-scaling costs a little more but prevents a class of 2 AM incidents entirely.
RDS Encryption and Security
Encryption at rest:
Must be enabled at instance creation — cannot be added afterUses AWS KMS for key managementEncrypts: data volumes, automated backups, snapshots, Read ReplicasHow to encrypt an existing unencrypted RDS instance (4 steps):
1. Take a snapshot of the unencrypted instance2. Copy the snapshot — enable encryption during the copy3. Restore the encrypted snapshot as a new DB instance4. Update application connection string to point to new instance5. Delete the old unencrypted instanceEncryption in transit:
SSL/TLS for all connectionsForce SSL: set rds.force_ssl = 1 in the DB Parameter GroupIAM Authentication — no passwords stored in app code:
Instead of username and password, your application generates a short-lived IAM authentication token and uses that to connect. The token is valid for 15 minutes. Works for MySQL and PostgreSQL.
Application calls AWS API → gets 15-minute auth tokenConnects to RDS using the token as the passwordNo DB password stored in your application config anywhereAmazon Aurora — Built Differently
Aurora is AWS's own relational database engine. It is PostgreSQL and MySQL compatible — your existing drivers, queries, and tools work without changes. But the internals are built from scratch for the cloud.
Aurora vs standard RDS:
| Feature | Standard RDS | Aurora |
|---|---|---|
| Performance vs MySQL | Baseline | 5x faster |
| Performance vs PostgreSQL | Baseline | 3x faster |
| Read Replica lag | Seconds | Under 10ms |
| Failover time | 60-120 seconds | Under 30 seconds |
| Storage growth | Manual — you choose size | Auto-grows 10 GB at a time, up to 128 TB |
| Storage copies | 2 (primary + standby) | 6 copies across 3 AZs automatically |
| Max Read Replicas | 5 | 15 |
| Cost vs standard RDS | Baseline | About 20% more |
Aurora storage architecture:
Aurora stores 6 copies of your data across 3 Availability Zones automatically. You do not configure this. It just happens.
ap-south-1a → Copy 1, Copy 2ap-south-1b → Copy 3, Copy 4ap-south-1c → Copy 5, Copy 6 Can lose 2 AZs and still readCan lose 1 AZ and still writeSelf-healing — corrupted blocks detected and replaced automaticallyWriter Endpoint and Reader Endpoint:
Aurora gives you two DNS endpoints:
Writer Endpoint → always points to the current primary (write here)Reader Endpoint → load balances across all Read Replicas (read here)If the primary fails and a replica is promoted, the Writer Endpoint automatically updates. Your application does not need to change anything.
RememberAurora automatically maintains 6 copies across 3 AZs with no configuration. Standard RDS Multi-AZ gives you 2 copies. Aurora gives 6 copies by default. This is why Aurora recovers faster and loses less data when things go wrong.
Aurora Serverless — Scale to Zero
For unpredictable or intermittent workloads, Aurora Serverless removes the concept of provisioned instances entirely.
Development database used 2 hours per day: Standard RDS → pays for 24/7 instance even when idle overnight Aurora Serverless → scales down to zero at night → pays for 2 hours of actual use Flash sale at Swiggy: Standard RDS → must pre-provision for peak traffic Aurora Serverless → scales up instantly when load arrives, scales back down afterYou define minimum and maximum Aurora Capacity Units. Aurora scales between them automatically based on load. Pay per second for actual compute consumed.
Best for: development databases, intermittent batch workloads, new applications where traffic is unknown.
Aurora Global Database — Cross-Region in Under 1 Second
Aurora Global creates up to 10 secondary read-only regions. Replication lag is under 1 second. A secondary region can be promoted to primary in under 1 minute.
Primary Region: ap-south-1 (Mumbai) All writes happen here Replication lag: 0 Secondary Region 1: ap-southeast-1 (Singapore) Read traffic from Southeast Asia served locally Replication lag: under 1 second Secondary Region 2: us-east-1 (Virginia) Read traffic from US served locally Replication lag: under 1 second Disaster Recovery: Mumbai becomes unavailable Promote Singapore to primary in under 1 minute All traffic shifts to Singapore Data loss: under 1 second of writes (RPO)Compared to standard RDS cross-region Read Replica:
| Aurora Global | RDS Cross-Region Read Replica | |
|---|---|---|
| Replication lag | Under 1 second | Minutes |
| Promote to primary | Under 1 minute | Minutes to hours |
| RPO | Under 1 second | Minutes |
Aurora Database Cloning
Create a copy of your production Aurora cluster for staging or testing in minutes — not hours.
Production Aurora cluster with 500 GB of data ↓Create clone (takes minutes, not hours) ↓Staging cluster with identical 500 GB of data appears ↓No actual data copied — uses copy-on-write internallyProduction cluster completely unaffectedCopy-on-write means the clone shares the same underlying storage as production. Only when you write to the clone does it start storing separate data for those changed blocks. Fast, cheap, and completely isolated from production.
Use cases:
- Test a schema migration against real production-sized data before running on prod
- Debug a production data issue without touching live data
- Run load tests against realistic dataset
RDS Backup Strategies
Automated Backups — Point in Time Recovery (PITR):
AWS continuously records transaction logsAutomated backups taken daily during your backup windowRetention: 1 to 35 days (set to 35 days for production)Restore: to any specific second within the retention windowRestore creates a NEW database instance — original is untouchedRememberSetting backup retention to 0 disables automated backups entirely and disables PITR. Always keep retention at minimum 7 days for any production database.
Manual Snapshots:
Taken on demand by you at any timeNever expire automatically — kept until you explicitly delete themCan be copied to another region for cross-region backupCan be shared with another AWS accountTake a manual snapshot before any risky operation (schema migration, major upgrade)The golden rule: Automated backups give you flexibility (restore to any second). Manual snapshots give you permanence (kept as long as you want). Use both.
ElastiCache as a Caching Layer
RDS can handle significant load, but every database has a ceiling. For read-heavy workloads where the same queries run thousands of times per minute — fetching a product page, getting a user profile, loading configuration — ElastiCache Redis sits in front of RDS and serves cached results in sub-millisecond time.
Without ElastiCache: 1000 users request the same product page simultaneously → 1000 identical SELECT queries hit RDS → RDS CPU climbs, latency increases for everyone With ElastiCache Redis: First request → cache miss → query RDS → store result in Redis Next 999 requests → cache hit → returned from memory in < 1ms RDS receives 1 query instead of 1000RememberElastiCache does not work transparently. Your application code must check the cache first, then fall back to RDS on a miss. This requires code changes — it is not a drop-in addition. Plan for it at design time.
Two caching patterns:
Lazy Loading: Check cache first Hit → return immediately (fast) Miss → query RDS → store in cache → return to user Only caches what is actually requested Session Store: User session stored in Redis instead of in EC2 memory Any instance behind the load balancer can serve any user Eliminates the need for sticky sessions entirelyHands-on Lab — Launch RDS, Connect from EC2, Create Read Replica
Step 1 — Create a Security Group for RDS
EC2 → Security Groups → Create Security GroupName: rds-sgVPC: your VPC Inbound rule:Type: MySQL/Aurora Port: 3306Source: App EC2 Security Group ID (not 0.0.0.0/0 — never expose DB to internet)Step 2 — Create a DB Subnet Group
RDS → Subnet Groups → Create DB Subnet GroupName: prod-db-subnet-groupVPC: your VPCAdd subnets: select your two private subnets in different AZsCreateStep 3 — Launch the RDS Instance
RDS → Databases → Create databaseEngine: MySQL Version: 8.0Template: Production Settings:DB identifier: devops-prod-dbMaster username: adminMaster password: set a strong password Instance: db.t3.micro (for lab — use db.t3.medium+ for real workloads)Storage: 20 GB gp3 Storage autoscaling: Enable Maximum: 100 GB Connectivity:VPC: your VPCSubnet group: prod-db-subnet-groupPublic access: No (always No for production)Security Group: rds-sg Additional:Backup retention: 7 daysEnable encryption: YesEnable Enhanced monitoring: Yes Create database — takes 3-5 minutesStep 4 — Get the endpoint
RDS → Databases → devops-prod-db → Connectivity & securityCopy the Endpoint addressIt looks like: devops-prod-db.abc123.ap-south-1.rds.amazonaws.comStep 5 — Connect from EC2
SSH into your EC2 instance in the same VPC## Install MySQL clientsudo yum install -y mysql ## Connect to RDS (replace with your endpoint and password)mysql -h devops-prod-db.abc123.ap-south-1.rds.amazonaws.com \ -u admin -p ## Create a test database and tableCREATE DATABASE devopsdb;USE devopsdb; CREATE TABLE engineers ( id INT AUTO_INCREMENT PRIMARY KEY, name VARCHAR(50), city VARCHAR(50)); INSERT INTO engineers (name, city) VALUES ('Rahul Sharma', 'Mumbai'), ('Priya Nair', 'Bengaluru'), ('Arjun Verma', 'Hyderabad'); SELECT * FROM engineers;Step 6 — Create a Read Replica
RDS → Databases → devops-prod-db → Actions → Create read replicaDB identifier: devops-prod-db-replicaRegion: same region (or another region for cross-region DR)Instance class: db.t3.microCreate read replica — takes 3-5 minutes After creation:Connect to the replica endpoint and run SELECT queriesVerify data is present — replication is workingStep 7 — Take a manual snapshot
RDS → Databases → devops-prod-db → Actions → Take snapshotName: devops-prod-db-manual-backup-before-migrationTake snapshot This gives you a permanent restore point before any risky operation.Step 8 — Cleanup
Delete Read Replica first (Actions → Delete)Delete Primary instance (Actions → Delete)Check: Create final snapshot? → No (for lab cleanup)Confirm deletionDelete DB Subnet GroupDelete Security GroupCommon Mistakes to Avoid
Common MistakeConfusing Read Replicas with Multi-AZ. A Read Replica serves read traffic and uses ASYNC replication — there can be lag, and it will NOT automatically take over if the primary fails. Multi-AZ standby serves zero traffic, uses SYNC replication, and takes over automatically when the primary fails. These are different features solving different problems. Both should be enabled in production.
Common MistakeRunning heavy reporting queries against the production primary. A slow GROUP BY or analytical query running 30 seconds holds locks and slows down every other query on the same instance. Point all reporting tools and dashboards to a Read Replica endpoint. The primary should only serve live application traffic.
Common MistakeNot enabling Storage Auto Scaling. Storage full is not a warning — it is an immediate write failure. Every insert and update starts failing. Enable Auto Scaling with a sensible maximum and this failure mode goes away permanently.
SecurityNever set Public Access to Yes on an RDS instance. Public access means the instance gets a public IP and becomes reachable from the internet. Even with a strong password, a public database is exposed to brute force attacks, credential stuffing, and any future engine CVE. Always put RDS in a private subnet.
TipAurora Database Cloning is one of the most underused features in RDS. Instead of restoring a snapshot (which copies all data and takes hours for large databases), cloning uses copy-on-write and takes minutes regardless of database size. Any team that regularly tests schema migrations should be using it.