Skip to main content

Multi-AZ RDS

An RDS high availability configuration that keeps a synchronous standby copy in a different Availability Zone. If the primary fails, RDS automatically promotes the standby in 60-120 seconds with no data loss.

What is Multi-AZ

Multi-AZ keeps your database alive when an Availability Zone, hardware, or software failure occurs. AWS maintains an exact synchronous copy in a different AZ — always identical, always ready.

◈ DIAGRAM
Normal state:
ALL traffic → Primary in ap-south-1a
Standby in ap-south-1b → receives SYNC replication, serves NOTHING
Primary fails (hardware, network, or software issue):
AWS detects failure
Standby promoted to primary automatically
DNS endpoint updated (same endpoint, new target)
Application reconnects → back online
Total downtime: 60-120 seconds
Data loss: zero (SYNC replication means standby was always identical)

SYNC Replication

TEXT
Primary writes a transaction
Primary waits for standby to confirm it received the write
Only then does the write complete
Result: standby is ALWAYS identical to primary — zero lag
This differs from Read Replica replication (ASYNC) which can have seconds of lag.

Multi-AZ Standby Does Not Serve Traffic

This is the most common misconception. The standby exists only for failover — it handles zero queries during normal operations.

TEXT
Wrong assumption: Multi-AZ doubles my read capacity
Reality: Multi-AZ only gives you automatic failover
To scale reads: create Read Replicas
To survive failures: enable Multi-AZ
In production: do both

Multi-AZ Cluster (new option)

Beyond the traditional 1 primary + 1 standby, Multi-AZ Cluster provides 1 writer + 2 readable standby instances across 3 AZs. Standbys can serve read traffic and automatic failover completes in under 35 seconds.

Failover Triggers

TEXT
Primary instance failure
AZ failure or outage
OS-level software issue on primary
Scheduled maintenance requiring a reboot
You manually trigger it: Reboot with Failover option
Remember

Enable Multi-AZ on every production RDS instance. The additional cost (you pay for the standby instance) is worth the automatic recovery from failures. Without Multi-AZ, a hardware failure means manual intervention and potentially hours of downtime.

Frequently Asked Questions

How is Multi-AZ RDS different from a read replica?

Multi-AZ maintains a synchronous standby in a different Availability Zone purely for failover — you cannot query the standby directly, and its only job is to take over automatically if the primary fails, typically within 60-120 seconds with zero data loss because replication is synchronous. A read replica, by contrast, is asynchronously replicated, is queryable for offloading read traffic, and does not automatically fail over — promoting one to primary is a manual (or scripted) action.

What's a common mistake teams make assuming Multi-AZ RDS protects against?

Multi-AZ protects against infrastructure failure — an AZ outage, hardware fault, storage failure — but it replicates data changes synchronously, so it does nothing to protect against a bad `DELETE`, a dropped table, or application-level corruption; that mistake gets replicated to the standby just as fast as any legitimate write. Point-in-time recovery from automated backups, not Multi-AZ, is the actual defense against those scenarios.