What you will learn
- How DNS resolution actually works step by step — Root, TLD, Authoritative Name Server
- Hosted zones — Public vs Private and what the $0.50/month buys you
- Record types — A, AAAA, CNAME, NS, and the Alias record that replaces CNAME for AWS resources
- CNAME vs Alias — the distinction that trips up every engineer at least once
- TTL strategy — how to change DNS records safely without locking clients to stale values
- All 7 routing policies and exactly when to use each one
- Health checks — endpoint monitoring, calculated health checks, and private resource patterns
- How to use Route 53 with a domain registered on GoDaddy or another registrar
- Hybrid DNS — Inbound and Outbound Resolver Endpoints for on-premises integration
Why this matters
At Zerodha, if the primary trading API region goes down at market open, Route 53 Failover routing detects the health check failure and automatically redirects traffic to the backup region — without any human intervention. At Swiggy, Latency-Based routing ensures users in Mumbai connect to ap-south-1 and users in Singapore connect to ap-southeast-1, getting the fastest response from the nearest region. At Hotstar, Weighted routing lets the team gradually shift 10% of traffic to a new version during a deployment and roll back instantly if error rates climb. These are not theoretical use cases — they are production patterns running at scale right now.
How DNS Resolution Works
Every time you type a domain name, a multi-step resolution process runs in milliseconds.
You type: api.swiggy.com in your browser Step 1: Browser checks local cache → not foundStep 2: Asks Local DNS Resolver (your ISP or 8.8.8.8)Step 3: Local Resolver asks Root DNS Server → "I don't know api.swiggy.com, try .com TLD server"Step 4: Local Resolver asks .com TLD Server → "I don't know, try swiggy.com Name Server"Step 5: Local Resolver asks swiggy.com Name Server (authoritative) → "api.swiggy.com is at 13.235.45.67"Step 6: Local Resolver returns IP to browser, caches it for the TTLStep 7: Browser connects to 13.235.45.67The whole process takes 20-100ms. Results are cached so subsequent requests skip most steps.
DNS terminology you must know:
| Term | What it means |
|---|---|
| Domain Registrar | Where you buy domain names — Route 53, GoDaddy, Namecheap |
| DNS Records | Instructions stored in DNS — A, AAAA, CNAME, NS, etc |
| Zone File | The file containing all DNS records for a domain |
| Authoritative Name Server | The server that has the definitive answer for your domain |
| TLD | Top Level Domain — .com, .in, .org, .gov |
| SLD | Second Level Domain — amazon.com, google.com |
| FQDN | Fully Qualified Domain Name — api.www.example.com |
What is Amazon Route 53
Route 53 is AWS's highly available, fully authoritative DNS service. Authoritative means you — the customer — can update DNS records directly. You are in full control of where traffic goes.
Route 53 is also a Domain Registrar. You can buy and register domain names directly through it, just like GoDaddy or Namecheap.
Client types: myapp.in → Route 53 receives the DNS query → Looks up the record for myapp.in → Returns: 54.22.33.44 → Client connects to EC2 at 54.22.33.44Key facts:
The only AWS service with a 100% availability SLA — AWS guarantees it never goes downCan check health of your resources and route based on healthThe name "Route 53" references port 53 — the standard DNS portCost: $0.50 per month per hosted zoneDNS Record Types
Each record is an instruction telling Route 53 how to respond to a query.
A Record — maps hostname to IPv4 address:
example.com → 1.2.3.4AAAA Record — maps hostname to IPv6 address:
example.com → 2001:0db8:85a3::8a2e:0370:7334CNAME Record — maps hostname to another hostname:
www.example.com → app.example.com Critical limitation: CANNOT create a CNAME for the root domain (Zone Apex) Works: www.example.com → anything.com (subdomain — fine) Fails: example.com → anything.com (root domain — not allowed)NS Record — Name Server records:
Tell the internet which DNS servers are authoritative for your domain. When you buy a domain on GoDaddy and point it to Route 53, you update NS records at GoDaddy to use Route 53's name servers.
Must know deeply: A, AAAA, CNAME, NS
Hosted Zones
A Hosted Zone is a container in Route 53 holding all DNS records for a domain and its subdomains.
Cost: $0.50 per month per hosted zone
Public Hosted Zone — routes traffic on the public internet:
Internet client → Public Hosted Zone (example.com) → EC2, ALB, CloudFrontAnyone on the internet can query a public hosted zone.
Private Hosted Zone — routes traffic only inside your VPC:
EC2 inside VPC → Private Hosted Zone (company.internal) db.company.internal → 10.0.0.35 (RDS instance private IP) cache.company.internal → 10.0.10.9 (ElastiCache private IP) api.company.internal → 10.0.5.20 (internal service)Only resources inside your VPC can resolve these names. Microservices communicate using friendly names instead of hardcoded private IPs.
TTL — Time To Live
TTL is how long a DNS resolver caches a record before asking Route 53 again.
High TTL (86400 seconds = 24 hours): Fewer Route 53 queries — cheaper Change a record → clients stuck with old value for up to 24 hours Good for stable records that rarely change Low TTL (60 seconds): More Route 53 queries — slightly more expensive Changes propagate within 60 seconds Good when actively making changes or migratingSafe strategy for changing a DNS record:
Step 1: Lower TTL to 60 secondsStep 2: Wait 24-48 hours (so all cached clients pick up the new low TTL)Step 3: Make your DNS changeStep 4: Wait 60 seconds (all clients get the new value)Step 5: Raise TTL back to 86400If you skip Step 2, some clients still have the old high TTL cached and will not see your DNS change for hours.
RememberTTL is mandatory on all records except Alias records. Alias records have no TTL — Route 53 manages it automatically.
CNAME vs Alias Records
AWS resources expose DNS hostnames, not static IPs:
Your Load Balancer DNS: lb1-1234.ap-south-1.elb.amazonaws.comYou want users to access: api.myapp.inYou need to point your domain at this AWS hostname. Two options:
CNAME:
Points a hostname to any other hostnameOnly works for non-root subdomains — www.example.com works, example.com does notCharges apply for DNS queries against CNAME recordsAlias Record:
Points a hostname specifically to an AWS resourceWorks for BOTH root and non-root domains — example.com and www.example.comFree — no query chargesAutomatically tracks IP changes of the AWS resource (ALB IPs rotate — Alias handles transparently)Built-in health check supportAlways type A or AAAA — never CNAME typeNo TTL — Route 53 manages it CNAME: www.example.com → lb.anything.com (works — non-root subdomain) example.com → lb.anything.com (fails — root domain, Zone Apex) Alias: www.example.com → lb.ap-south-1.elb.amazonaws.com (works) example.com → lb.ap-south-1.elb.amazonaws.com (works — root supported)Valid Alias record targets:
Elastic Load Balancers (ALB, NLB, CLB)CloudFront DistributionsAPI GatewayElastic Beanstalk environmentsS3 WebsitesVPC Interface EndpointsGlobal AcceleratorRoute 53 record in the same hosted zoneRememberYou cannot set an Alias record pointing to an EC2 DNS name. EC2 is not a valid Alias target. Only the services listed above are supported.
Health Checks
Route 53 Health Checks monitor your resources and trigger automatic DNS failover when something goes wrong. Health checks work only for public resources — they cannot directly access private VPC resources.
Type 1 — Monitoring an Endpoint:
Route 53 sends health checker requests from 15 global locations to your endpoint.
Protocol: HTTP, HTTPS, or TCPInterval: 30 seconds (default) or 10 seconds (higher cost)Healthy threshold: 3 consecutive successesHealthy if: 18% or more of health checkers report healthyHealth check passes when: endpoint returns 2xx or 3xx status codeCan also check: specific text in the first 5,120 bytes of the responseType 2 — Calculated Health Checks:
Combine multiple child health checks into one parent using AND/OR/NOT logic.
Child health check A (region 1) ─┐Child health check B (region 2) ─┼──> Parent health check (AND logic)Child health check C (region 3) ─┘ Specify how many children must pass for parent to be healthyUse: doing maintenance on one resource without failing overall healthType 3 — Health Checks for Private Resources:
Route 53 health checkers run outside your VPC — they cannot reach private endpoints directly.
Workaround: EC2 in private subnet ↓ CloudWatch Metric (monitors CPU, errors, or custom metric) ↓ CloudWatch Alarm (fires when threshold crossed) ↓ Route 53 Health Check (monitors the CloudWatch Alarm state)Route 53 indirectly monitors private resources through the CloudWatch Alarm state.
The 7 Routing Policies
1. Simple Routing — single resource, no health checks:
Can return multiple IP values in the same recordIf multiple values returned → client picks one at randomCannot be associated with Health Checks Client: "What is api.example.com?"Route 53 returns: 11.22.33.44, 55.66.77.88, 99.11.22.33Client picks one at random and connects2. Weighted Routing — control traffic percentage:
Assign relative weights to each recordTraffic % = record weight / sum of all weightsCan be associated with Health ChecksWeight = 0 → stop sending traffic to that resource api.example.com → 11.22.33.44 (Weight: 70) → 70% of trafficapi.example.com → 55.66.77.88 (Weight: 20) → 20% of trafficapi.example.com → 99.11.22.33 (Weight: 10) → 10% of traffic Use for: A/B testing new versions, gradual blue-green deployments, distributing load across regions3. Latency-Based Routing — route to lowest latency region:
Latency measured between users and AWS Regions — not geographic distanceA user in Germany may be routed to us-east-1 if it has lower latency than eu-west-1Can be associated with Health Checks User in Mumbai: Route 53 checks latency: ap-south-1=12ms, us-east-1=180ms, eu-west-1=210ms Returns: ap-south-1 (lowest latency wins)4. Failover Routing — active-passive disaster recovery:
Health Check is mandatory on the primary recordIf primary health check fails → Route 53 automatically returns secondary record Normal: Client → Route 53 → Primary EC2 (healthy)Failure: Primary health check fails → Route 53 automatically returns Secondary EC2 → Client connects to Secondary (no manual action)5. Geolocation Routing — route by physical location:
Route based on where the user physically is — not latencySpecify routing by Continent, Country, or US StateMost specific location wins if rules overlapAlways create a Default record for users whose location matches nothing User in Germany → EU server (German language content)User in India → AP server (regional pricing and content)User anywhere else → Default record (global fallback) Use for: website localisation, content restriction by country, regional complianceRememberGeolocation routes by where the user physically is. Latency-based routes by which region responds fastest. A user physically close to a region may still have higher latency — they are different things.
6. Geoproximity Routing — shift traffic by adjusting bias:
Route based on geographic location of users AND resourcesBias value shifts more or less traffic toward a specific resource Positive bias (+1 to +99): expands coverage, attracts more traffic Negative bias (-1 to -99): shrinks coverage, attracts less trafficRequires Route 53 Traffic Flow feature Scenario: gradually shift traffic from us-west-1 to us-east-1 Both at bias 0 → traffic split by proximity us-east-1 bias = 50 → coverage expands → more traffic goes there us-east-1 bias = 99 → almost all traffic shifted to us-east-17. IP-Based Routing — route by client IP address:
You provide a list of CIDR ranges mapped to specific endpointsRoute 53 routes based on the client's IP address User A (IP: 203.0.113.x) → matches CIDR 203.0.113.0/24 → EC2 in Region AUser B (IP: 200.5.4.x) → matches CIDR 200.5.4.0/24 → EC2 in Region B Use for: route known ISP IP ranges to nearby endpoints, reduce data transfer costs for specific networks8. Multi-Value Routing — healthy IPs only:
Returns up to 8 healthy records per queryEach record associated with a Health Check — only healthy includedClient receives multiple IPs and picks oneNot a substitute for a Load Balancer — client-side selection only Client: "What is api.example.com?"Route 53 checks health of all records: 192.0.2.2 → HEALTHY ← returned 198.51.100.2 → HEALTHY ← returned 203.0.113.2 → UNHEALTHY ← excluded Returns only healthy IPs. Client picks one.Using GoDaddy as Registrar with Route 53 for DNS
You can buy a domain on GoDaddy and manage DNS records on Route 53.
1. Buy example.in on GoDaddy2. Create a Public Hosted Zone in Route 53 for example.in3. Route 53 gives you 4 Name Server (NS) records4. Go to GoDaddy → Manage Domain → Change Nameservers5. Paste Route 53's 4 NS records into GoDaddy6. GoDaddy tells the internet: "Route 53 manages DNS for example.in"7. All DNS queries for example.in now go to Route 53From this point manage all DNS records inside Route 53 even though you bought the domain on GoDaddy.
Hybrid DNS — Resolver Endpoints
By default, EC2 instances inside your VPC can resolve AWS service DNS names and private hosted zone names. But when you have an on-premises data center connected via VPN or Direct Connect, DNS needs to work in both directions.
Inbound Endpoint — on-premises to AWS:
On-premises server needs to resolve: db.aws.internal → On-premises DNS Resolver sends query to Route 53 Inbound Endpoint → Route 53 looks up the Private Hosted Zone → Returns: db.aws.internal → 10.0.1.5 → On-premises server connects to RDS in your VPCOutbound Endpoint — AWS to on-premises:
EC2 in VPC needs to resolve: erp.onpremise.internal → Route 53 Resolver: no match in any hosted zone → Forwards query to Outbound Endpoint → Outbound Endpoint queries on-premises DNS Resolvers via VPN/Direct Connect → On-premises DNS returns: erp.onpremise.internal → 192.168.1.10 → EC2 connects to the on-premises ERP serverBoth endpoints operate over VPN or Direct Connect to bridge DNS between AWS and on-premises.
Hands-on Lab — Hosted Zone, Records, and Weighted Routing
Step 1 — Create a Public Hosted Zone
Route 53 → Hosted zones → Create hosted zoneDomain name: devops-lab.in (or any domain you own)Type: Public hosted zoneCreate hosted zone Note the 4 NS records — update these at your domain registrarto point DNS to Route 53.Step 2 — Create a simple A record
Route 53 → Hosted zones → devops-lab.in → Create recordRecord name: wwwRecord type: AValue: your EC2 public IPTTL: 60 (low for testing — change to 300+ in production)Create records Test: dig www.devops-lab.in @8.8.8.8Should return your EC2 IP within 60 seconds.Step 3 — Create a Health Check
Route 53 → Health checks → Create health checkName: server-1-healthWhat to monitor: EndpointProtocol: HTTP Domain: your EC2 public IP Port: 80 Path: /Request interval: 30 seconds Failure threshold: 3Create health check Wait 2 minutes → Health check status should show Healthy.Step 4 — Create Weighted routing records
Route 53 → Hosted zones → devops-lab.in → Create recordRecord name: apiRecord type: ARouting policy: WeightedValue: INSTANCE-1-IP Weight: 70 Record ID: server-1Health check: server-1-healthAdd another record:Value: INSTANCE-2-IP Weight: 30 Record ID: server-2Create recordsStep 5 — Test traffic distribution
Run multiple DNS lookups and observe which IP returns:For 10 queries, approximately 7 should return Server 1 IP,3 should return Server 2 IP. dig api.devops-lab.in @8.8.8.8(run several times — watch the answer change between the two IPs)Step 6 — Test health check failover
Stop the EC2 instance behind Server 1.Wait 90 seconds for health check to detect failure (3 checks × 30s). Route 53 → Health checks → server-1-health → Status: Unhealthy Now all 10/10 DNS queries return Server 2 IP.Route 53 automatically excluded the unhealthy endpoint.Step 7 — Cleanup
Route 53 → Hosted zones → devops-lab.in → select all records except NS and SOADelete recordsRoute 53 → Health checks → server-1-health → DeleteRoute 53 → Hosted zones → devops-lab.in → Delete hosted zoneProduction Best Practices and Common Pitfalls
- Always use Alias records instead of CNAME for AWS resources — free, supports root domain, auto-tracks IP changes
- Lower TTL before making DNS changes — if TTL is 24 hours and you change a record, clients are stuck for up to 24 hours
- Always create a Default record with Geolocation routing — without it, users from unmatched locations get NXDOMAIN errors
- Use health checks with Failover routing — Failover without health checks means Route 53 keeps routing to a dead endpoint
- Use Latency-Based routing for multi-region deployments — it automatically finds the fastest region for each user
- Do not confuse Geolocation with Latency-Based — Geolocation routes by where users are, Latency routes by where response is fastest
Quick Reference and Troubleshooting Commands
| Task | Command |
|---|---|
| List hosted zones | aws route53 list-hosted-zones |
| Create hosted zone | aws route53 create-hosted-zone --name example.com --caller-reference $(date +%s) |
| List records in zone | aws route53 list-resource-record-sets --hosted-zone-id <ZONE-ID> |
| Create/update record | aws route53 change-resource-record-sets --hosted-zone-id <ZONE-ID> --change-batch file://record.json |
| List health checks | aws route53 list-health-checks |
| Get health check status | aws route53 get-health-check-status --health-check-id <ID> |
| Delete health check | aws route53 delete-health-check --health-check-id <ID> |
| Check DNS resolution | dig api.example.com @8.8.8.8 |
| Check which NS records serve a domain | dig NS example.com |
Common problems and fixes:
| Problem | Likely cause | Fix |
|---|---|---|
| DNS change not propagating | Old TTL still cached by resolvers | Wait for old TTL to expire or lower TTL before future changes |
| NXDOMAIN for geolocation users | No Default record created | Add a Default record as fallback for unmatched locations |
| Failover not working | Health check not associated with primary record | Associate health check with primary record in the record set |
| Cannot create CNAME for root domain | DNS standard limitation | Use Alias record instead — supports root domain |
| Private hosted zone not resolving | VPC not associated with the hosted zone | Associate the VPC with the private hosted zone |
Common MistakeCreating a CNAME for the root domain (example.com). The DNS standard does not allow CNAME at the Zone Apex. Route 53 will reject it. Use an Alias record — it achieves the same result, works for root domains, and is free.
Common MistakeUsing Failover routing without health checks. Without a health check, Route 53 has no way to know the primary is down and will keep routing to a failed endpoint indefinitely. Health checks are mandatory for Failover routing to actually fail over.
TipFor a production multi-region setup combine Latency-Based routing with health checks on every record. Latency-Based sends users to the fastest region normally. Health checks automatically exclude unhealthy regions. If ap-south-1 goes down during IPL streaming, Route 53 automatically routes users to the next fastest healthy region — no manual intervention, no ops team paged at 3 AM.
Common Mistakes to Avoid
Common MistakeUsing a CNAME for the root domain (example.com). DNS does not allow CNAME at the Zone Apex. Use an Alias record instead — it supports root domains, is free, and automatically tracks IP changes of the AWS resource.
Common MistakeUsing Failover routing without a health check on the primary record. Without a health check, Route 53 has no way to know the primary is down. It keeps routing to the failed endpoint indefinitely. Health checks are mandatory for Failover routing to actually fail over.
TipLower your TTL to 60 seconds before making any DNS change. If TTL is 24 hours and you change a record, clients with the old value cached are stuck for up to 24 hours. Lower TTL → wait 24 hours for caches to expire → make the change → raise TTL back.