Multi-AZ RDS
An RDS high availability configuration that keeps a synchronous standby copy in a different Availability Zone. If the primary fails, RDS automatically promotes the standby in 60-120 seconds with no data loss.
What is Multi-AZ
Multi-AZ keeps your database alive when an Availability Zone, hardware, or software failure occurs. AWS maintains an exact synchronous copy in a different AZ — always identical, always ready.
Normal state: ALL traffic → Primary in ap-south-1a Standby in ap-south-1b → receives SYNC replication, serves NOTHING Primary fails (hardware, network, or software issue): AWS detects failure Standby promoted to primary automatically DNS endpoint updated (same endpoint, new target) Application reconnects → back online Total downtime: 60-120 seconds Data loss: zero (SYNC replication means standby was always identical)SYNC Replication
Primary writes a transactionPrimary waits for standby to confirm it received the writeOnly then does the write completeResult: standby is ALWAYS identical to primary — zero lag This differs from Read Replica replication (ASYNC) which can have seconds of lag.Multi-AZ Standby Does Not Serve Traffic
This is the most common misconception. The standby exists only for failover — it handles zero queries during normal operations.
Wrong assumption: Multi-AZ doubles my read capacityReality: Multi-AZ only gives you automatic failover To scale reads: create Read ReplicasTo survive failures: enable Multi-AZIn production: do bothMulti-AZ Cluster (new option)
Beyond the traditional 1 primary + 1 standby, Multi-AZ Cluster provides 1 writer + 2 readable standby instances across 3 AZs. Standbys can serve read traffic and automatic failover completes in under 35 seconds.
Failover Triggers
Primary instance failureAZ failure or outageOS-level software issue on primaryScheduled maintenance requiring a rebootYou manually trigger it: Reboot with Failover optionRememberEnable Multi-AZ on every production RDS instance. The additional cost (you pay for the standby instance) is worth the automatic recovery from failures. Without Multi-AZ, a hardware failure means manual intervention and potentially hours of downtime.
Frequently Asked Questions
How is Multi-AZ RDS different from a read replica?
Multi-AZ maintains a synchronous standby in a different Availability Zone purely for failover — you cannot query the standby directly, and its only job is to take over automatically if the primary fails, typically within 60-120 seconds with zero data loss because replication is synchronous. A read replica, by contrast, is asynchronously replicated, is queryable for offloading read traffic, and does not automatically fail over — promoting one to primary is a manual (or scripted) action.
What's a common mistake teams make assuming Multi-AZ RDS protects against?
Multi-AZ protects against infrastructure failure — an AZ outage, hardware fault, storage failure — but it replicates data changes synchronously, so it does nothing to protect against a bad `DELETE`, a dropped table, or application-level corruption; that mistake gets replicated to the standby just as fast as any legitimate write. Point-in-time recovery from automated backups, not Multi-AZ, is the actual defense against those scenarios.