Skip to main content

Amazon RDS and Aurora - Managed Relational Databases in Production

Configure RDS Multi-AZ, Read Replicas, Aurora Global, and automated backups for production relational database workloads with zero-downtime operations.

What you will learn

  • What RDS is and which database engines it manages for you
  • Read Replicas vs Multi-AZ — the most important distinction in RDS and why they solve completely different problems
  • Storage Auto Scaling and why running out of disk causes immediate failures
  • How to encrypt RDS and authenticate with IAM instead of passwords
  • Aurora architecture — what makes it fundamentally different from standard RDS
  • Aurora Serverless for workloads that need to scale to zero
  • Aurora Global for sub-second cross-region replication
  • Aurora Database Cloning — production data in staging in minutes
  • RDS backup strategies — automated backups vs manual snapshots
  • ElastiCache as a caching layer to take load off RDS

Why this matters

Zerodha processes millions of stock transactions every day. Every trade, every ledger entry, every user position must be written atomically and consistently — this is what relational databases are built for. Razorpay's payment records need relational integrity across users, merchants, and transactions. Swiggy's order service uses Aurora because it needs the performance and reliability of a cloud-native database engine with the familiarity of MySQL.

Choosing the wrong database engine, forgetting to enable Multi-AZ, or not understanding the difference between Read Replicas and Multi-AZ are decisions that teams regret the moment their first production database incident hits.

What is Amazon RDS

RDS (Relational Database Service) is a managed service that runs standard SQL databases on AWS. The key word is managed — AWS handles the parts nobody wants to handle:

TEXT
AWS manages:
Provisioning the underlying EC2 and storage
OS patching and maintenance
Automated backups
Multi-AZ replication and failover
Minor version upgrades (optional)
Monitoring and CloudWatch metrics
You manage:
Schema design and queries
Application-level optimisation
Security Groups and network access
Major version upgrades
Who has database credentials

You cannot SSH into the underlying RDS server. That is the trade-off. You give up OS-level access and get a fully managed, highly available database in return.

Supported engines:

◈ DIAGRAM
PostgreSQL → most popular for new production workloads
MySQL → most widely used overall
MariaDB → open-source MySQL-compatible
Oracle → enterprise, bring-your-own-license
Microsoft SQL → Windows-centric enterprise
Aurora → AWS's own engine (covered in detail below)

Read Replicas vs Multi-AZ — The Most Important Distinction

This is the concept that every RDS question in production and in exams comes back to. Read Replicas and Multi-AZ solve completely different problems. They are often used together.

Read Replicas — for scaling reads

A Read Replica is a copy of your database that can serve read traffic. It uses asynchronous replication — changes flow from the primary to the replica, but there can be a small lag.

◈ DIAGRAM
Purpose: Scale out read traffic across multiple copies
Replication: ASYNC — replica is eventually consistent with primary
Who serves: Replica actively serves read queries from your application
Max replicas: 5 for standard RDS, 15 for Aurora
Failover: NOT automatic — you must manually promote a replica to primary
Without Read Replica:
All reads and writes → single RDS primary
Heavy reporting query running at 2 PM → slows down production writes
With Read Replica:
Writes + critical reads → RDS Primary
Reporting queries, analytics → Read Replica (no impact on primary)

Cross-region Read Replicas are also possible — useful when you want users in another region to read from a nearby database, or when you need disaster recovery in another region.

Multi-AZ — for disaster recovery

Multi-AZ keeps a standby copy of your database in a different Availability Zone using synchronous replication. The standby serves zero traffic during normal operations — it exists only to take over when the primary fails.

◈ DIAGRAM
Purpose: High availability and automatic failover
Replication: SYNC — standby is always identical to primary (no lag)
Who serves: Standby serves ZERO traffic during normal operations
Failover: Automatic — 60 to 120 seconds, no human action needed
Normal state:
All traffic (reads AND writes) → RDS Primary in AZ-1a
RDS Standby in AZ-1b → receives SYNC replication, does nothing else
Primary fails:
AWS detects failure
DNS endpoint automatically points to the standby
Your application reconnects to the new primary
Total downtime: 60 to 120 seconds
Zero data loss — SYNC replication means standby was always current
Remember

Read Replicas = read scaling via ASYNC replication. Multi-AZ = disaster recovery via SYNC replication. A Read Replica can have replication lag. Multi-AZ standby is always identical. Multi-AZ standby cannot serve read traffic. Read Replicas can. They solve different problems. Use both in production.

Can you promote a Read Replica to Multi-AZ?

Not directly. They are separate features. In production your setup typically looks like:

◈ DIAGRAM
RDS Primary (Multi-AZ enabled)
→ SYNC to RDS Standby in different AZ (serves no traffic)
→ ASYNC to Read Replica (serves read traffic, possibly in another region)

Storage Auto Scaling

Without Auto Scaling, when your RDS storage fills up, writes start failing immediately. Applications crash. Users see errors. Storage full is not a warning — it is an outage.

With Storage Auto Scaling, RDS detects when free space is running low and automatically expands storage — no downtime, no manual action.

TEXT
You set: Maximum Storage Threshold (e.g. 100 GB)
RDS monitors: free storage continuously
Trigger: free space < 10% AND low for 5 minutes AND 6 hours since last expansion
Action: storage expands automatically in increments
Tip

Enable Storage Auto Scaling on every production RDS instance from day one. Storage running out is an avoidable failure mode. The auto-scaling costs a little more but prevents a class of 2 AM incidents entirely.

RDS Encryption and Security

Encryption at rest:

TEXT
Must be enabled at instance creation — cannot be added after
Uses AWS KMS for key management
Encrypts: data volumes, automated backups, snapshots, Read Replicas

How to encrypt an existing unencrypted RDS instance (4 steps):

TEXT
1. Take a snapshot of the unencrypted instance
2. Copy the snapshot — enable encryption during the copy
3. Restore the encrypted snapshot as a new DB instance
4. Update application connection string to point to new instance
5. Delete the old unencrypted instance

Encryption in transit:

TEXT
SSL/TLS for all connections
Force SSL: set rds.force_ssl = 1 in the DB Parameter Group

IAM Authentication — no passwords stored in app code:

Instead of username and password, your application generates a short-lived IAM authentication token and uses that to connect. The token is valid for 15 minutes. Works for MySQL and PostgreSQL.

◈ DIAGRAM
Application calls AWS API → gets 15-minute auth token
Connects to RDS using the token as the password
No DB password stored in your application config anywhere

Amazon Aurora — Built Differently

Aurora is AWS's own relational database engine. It is PostgreSQL and MySQL compatible — your existing drivers, queries, and tools work without changes. But the internals are built from scratch for the cloud.

Aurora vs standard RDS:

Feature Standard RDS Aurora
Performance vs MySQL Baseline 5x faster
Performance vs PostgreSQL Baseline 3x faster
Read Replica lag Seconds Under 10ms
Failover time 60-120 seconds Under 30 seconds
Storage growth Manual — you choose size Auto-grows 10 GB at a time, up to 128 TB
Storage copies 2 (primary + standby) 6 copies across 3 AZs automatically
Max Read Replicas 5 15
Cost vs standard RDS Baseline About 20% more

Aurora storage architecture:

Aurora stores 6 copies of your data across 3 Availability Zones automatically. You do not configure this. It just happens.

◈ DIAGRAM
ap-south-1a → Copy 1, Copy 2
ap-south-1b → Copy 3, Copy 4
ap-south-1c → Copy 5, Copy 6
Can lose 2 AZs and still read
Can lose 1 AZ and still write
Self-healing — corrupted blocks detected and replaced automatically

Writer Endpoint and Reader Endpoint:

Aurora gives you two DNS endpoints:

◈ DIAGRAM
Writer Endpoint → always points to the current primary (write here)
Reader Endpoint → load balances across all Read Replicas (read here)

If the primary fails and a replica is promoted, the Writer Endpoint automatically updates. Your application does not need to change anything.

Remember

Aurora automatically maintains 6 copies across 3 AZs with no configuration. Standard RDS Multi-AZ gives you 2 copies. Aurora gives 6 copies by default. This is why Aurora recovers faster and loses less data when things go wrong.

Aurora Serverless — Scale to Zero

For unpredictable or intermittent workloads, Aurora Serverless removes the concept of provisioned instances entirely.

◈ DIAGRAM
Development database used 2 hours per day:
Standard RDS → pays for 24/7 instance even when idle overnight
Aurora Serverless → scales down to zero at night → pays for 2 hours of actual use
Flash sale at Swiggy:
Standard RDS → must pre-provision for peak traffic
Aurora Serverless → scales up instantly when load arrives, scales back down after

You define minimum and maximum Aurora Capacity Units. Aurora scales between them automatically based on load. Pay per second for actual compute consumed.

Best for: development databases, intermittent batch workloads, new applications where traffic is unknown.

Aurora Global Database — Cross-Region in Under 1 Second

Aurora Global creates up to 10 secondary read-only regions. Replication lag is under 1 second. A secondary region can be promoted to primary in under 1 minute.

TEXT
Primary Region: ap-south-1 (Mumbai)
All writes happen here
Replication lag: 0
Secondary Region 1: ap-southeast-1 (Singapore)
Read traffic from Southeast Asia served locally
Replication lag: under 1 second
Secondary Region 2: us-east-1 (Virginia)
Read traffic from US served locally
Replication lag: under 1 second
Disaster Recovery:
Mumbai becomes unavailable
Promote Singapore to primary in under 1 minute
All traffic shifts to Singapore
Data loss: under 1 second of writes (RPO)

Compared to standard RDS cross-region Read Replica:

Aurora Global RDS Cross-Region Read Replica
Replication lag Under 1 second Minutes
Promote to primary Under 1 minute Minutes to hours
RPO Under 1 second Minutes

Aurora Database Cloning

Create a copy of your production Aurora cluster for staging or testing in minutes — not hours.

◈ DIAGRAM
Production Aurora cluster with 500 GB of data
↓
Create clone (takes minutes, not hours)
↓
Staging cluster with identical 500 GB of data appears
↓
No actual data copied — uses copy-on-write internally
Production cluster completely unaffected

Copy-on-write means the clone shares the same underlying storage as production. Only when you write to the clone does it start storing separate data for those changed blocks. Fast, cheap, and completely isolated from production.

Use cases:

  • Test a schema migration against real production-sized data before running on prod
  • Debug a production data issue without touching live data
  • Run load tests against realistic dataset

RDS Backup Strategies

Automated Backups — Point in Time Recovery (PITR):

TEXT
AWS continuously records transaction logs
Automated backups taken daily during your backup window
Retention: 1 to 35 days (set to 35 days for production)
Restore: to any specific second within the retention window
Restore creates a NEW database instance — original is untouched
Remember

Setting backup retention to 0 disables automated backups entirely and disables PITR. Always keep retention at minimum 7 days for any production database.

Manual Snapshots:

TEXT
Taken on demand by you at any time
Never expire automatically — kept until you explicitly delete them
Can be copied to another region for cross-region backup
Can be shared with another AWS account
Take a manual snapshot before any risky operation (schema migration, major upgrade)

The golden rule: Automated backups give you flexibility (restore to any second). Manual snapshots give you permanence (kept as long as you want). Use both.

ElastiCache as a Caching Layer

RDS can handle significant load, but every database has a ceiling. For read-heavy workloads where the same queries run thousands of times per minute — fetching a product page, getting a user profile, loading configuration — ElastiCache Redis sits in front of RDS and serves cached results in sub-millisecond time.

◈ DIAGRAM
Without ElastiCache:
1000 users request the same product page simultaneously
→ 1000 identical SELECT queries hit RDS
→ RDS CPU climbs, latency increases for everyone
With ElastiCache Redis:
First request → cache miss → query RDS → store result in Redis
Next 999 requests → cache hit → returned from memory in < 1ms
RDS receives 1 query instead of 1000
Remember

ElastiCache does not work transparently. Your application code must check the cache first, then fall back to RDS on a miss. This requires code changes — it is not a drop-in addition. Plan for it at design time.

Two caching patterns:

◈ DIAGRAM
Lazy Loading:
Check cache first
Hit → return immediately (fast)
Miss → query RDS → store in cache → return to user
Only caches what is actually requested
Session Store:
User session stored in Redis instead of in EC2 memory
Any instance behind the load balancer can serve any user
Eliminates the need for sticky sessions entirely

Hands-on Lab — Launch RDS, Connect from EC2, Create Read Replica

Step 1 — Create a Security Group for RDS

◈ DIAGRAM
EC2 → Security Groups → Create Security Group
Name: rds-sg
VPC: your VPC
Inbound rule:
Type: MySQL/Aurora Port: 3306
Source: App EC2 Security Group ID (not 0.0.0.0/0 — never expose DB to internet)

Step 2 — Create a DB Subnet Group

◈ DIAGRAM
RDS → Subnet Groups → Create DB Subnet Group
Name: prod-db-subnet-group
VPC: your VPC
Add subnets: select your two private subnets in different AZs
Create

Step 3 — Launch the RDS Instance

◈ DIAGRAM
RDS → Databases → Create database
Engine: MySQL Version: 8.0
Template: Production
Settings:
DB identifier: devops-prod-db
Master username: admin
Master password: set a strong password
Instance: db.t3.micro (for lab — use db.t3.medium+ for real workloads)
Storage: 20 GB gp3 Storage autoscaling: Enable Maximum: 100 GB
Connectivity:
VPC: your VPC
Subnet group: prod-db-subnet-group
Public access: No (always No for production)
Security Group: rds-sg
Additional:
Backup retention: 7 days
Enable encryption: Yes
Enable Enhanced monitoring: Yes
Create database — takes 3-5 minutes

Step 4 — Get the endpoint

◈ DIAGRAM
RDS → Databases → devops-prod-db → Connectivity & security
Copy the Endpoint address
It looks like: devops-prod-db.abc123.ap-south-1.rds.amazonaws.com

Step 5 — Connect from EC2

TEXT
SSH into your EC2 instance in the same VPC
Bash
## Install MySQL client
sudo yum install -y mysql
## Connect to RDS (replace with your endpoint and password)
mysql -h devops-prod-db.abc123.ap-south-1.rds.amazonaws.com \
-u admin -p
## Create a test database and table
CREATE DATABASE devopsdb;
USE devopsdb;
CREATE TABLE engineers (
id INT AUTO_INCREMENT PRIMARY KEY,
name VARCHAR(50),
city VARCHAR(50)
);
INSERT INTO engineers (name, city) VALUES
('Rahul Sharma', 'Mumbai'),
('Priya Nair', 'Bengaluru'),
('Arjun Verma', 'Hyderabad');
SELECT * FROM engineers;

Step 6 — Create a Read Replica

◈ DIAGRAM
RDS → Databases → devops-prod-db → Actions → Create read replica
DB identifier: devops-prod-db-replica
Region: same region (or another region for cross-region DR)
Instance class: db.t3.micro
Create read replica — takes 3-5 minutes
After creation:
Connect to the replica endpoint and run SELECT queries
Verify data is present — replication is working

Step 7 — Take a manual snapshot

◈ DIAGRAM
RDS → Databases → devops-prod-db → Actions → Take snapshot
Name: devops-prod-db-manual-backup-before-migration
Take snapshot
This gives you a permanent restore point before any risky operation.

Step 8 — Cleanup

◈ DIAGRAM
Delete Read Replica first (Actions → Delete)
Delete Primary instance (Actions → Delete)
Check: Create final snapshot? → No (for lab cleanup)
Confirm deletion
Delete DB Subnet Group
Delete Security Group

Common Mistakes to Avoid

Common Mistake

Confusing Read Replicas with Multi-AZ. A Read Replica serves read traffic and uses ASYNC replication — there can be lag, and it will NOT automatically take over if the primary fails. Multi-AZ standby serves zero traffic, uses SYNC replication, and takes over automatically when the primary fails. These are different features solving different problems. Both should be enabled in production.

Common Mistake

Running heavy reporting queries against the production primary. A slow GROUP BY or analytical query running 30 seconds holds locks and slows down every other query on the same instance. Point all reporting tools and dashboards to a Read Replica endpoint. The primary should only serve live application traffic.

Common Mistake

Not enabling Storage Auto Scaling. Storage full is not a warning — it is an immediate write failure. Every insert and update starts failing. Enable Auto Scaling with a sensible maximum and this failure mode goes away permanently.

Security

Never set Public Access to Yes on an RDS instance. Public access means the instance gets a public IP and becomes reachable from the internet. Even with a strong password, a public database is exposed to brute force attacks, credential stuffing, and any future engine CVE. Always put RDS in a private subnet.

Tip

Aurora Database Cloning is one of the most underused features in RDS. Instead of restoring a snapshot (which copies all data and takes hours for large databases), cloning uses copy-on-write and takes minutes regardless of database size. Any team that regularly tests schema migrations should be using it.

Resources

AWS Direct Connect vs Site-to-Site VPN Failover

AWS Direct Connect vs Site-to-Site VPN Failover

Direct Connect vs VPN isn't really either/or for production — it's a primary-plus-failover pattern. Here's how to design it, and when either/or is right.

5 min read•Aug 2026
Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot, at the crossover where Lambda stops being cheaper — 2026 pricing, invocation thresholds, and interruption math.

5 min read•Aug 2026
Secrets Manager vs Parameter Store vs Vault

Secrets Manager vs Parameter Store vs Vault

AWS Secrets Manager, Parameter Store, and HashiCorp Vault compared for 2026 - cost math, rotation, multi-cloud fit, and the Vault-to-OpenBao fork.

5 min read•Aug 2026
AWS VPC Security: Hardening Every Layer

AWS VPC Security: Hardening Every Layer

Most cloud security incidents start with a misconfigured VPC. Here's how to harden every layer — subnets, Security Groups, NACLs, and IAM — for production.

5 min read•Jul 2026
Event-Driven Architecture on AWS Explained

Event-Driven Architecture on AWS Explained

Event-driven architecture on AWS decouples services and absorbs traffic spikes using SQS, SNS, EventBridge, and Lambda — workflows that scale themselves.

5 min read•Jul 2026
S3 vs RDS vs DynamoDB: Choosing AWS Storage

S3 vs RDS vs DynamoDB: Choosing AWS Storage

Choosing S3, RDS, or DynamoDB wrong costs you in performance, cost, and scalability. Here is a practical decision guide based on your actual access patterns.

5 min read•Jul 2026
AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS bills surprise teams every month. Here are the 8 concrete actions that cut cloud spend by 40-60% without touching your application architecture.

5 min read•Jul 2026
EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2, Lambda, or Fargate — choosing the wrong AWS compute option costs you money and performance. Here is exactly when to use each one in production.

5 min read•Jul 2026

Explore More in AWS Storage and Databases

All 6 Topics

Frequently Asked Questions

Is Amazon RDS and Aurora - Managed Relational Databases in Production free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Amazon RDS and Aurora - Managed Relational Databases in Production topic cover?

Configure RDS Multi-AZ, Read Replicas, Aurora Global, and automated backups for production relational database workloads with zero-downtime operations.