Skip to main content

Amazon Route 53 - DNS Routing Policies for Production Architectures

Configure Route 53 hosted zones, routing policies, health checks, and hybrid DNS endpoints to control global traffic routing and automatic failover.

What you will learn

  • How DNS resolution actually works step by step — Root, TLD, Authoritative Name Server
  • Hosted zones — Public vs Private and what the $0.50/month buys you
  • Record types — A, AAAA, CNAME, NS, and the Alias record that replaces CNAME for AWS resources
  • CNAME vs Alias — the distinction that trips up every engineer at least once
  • TTL strategy — how to change DNS records safely without locking clients to stale values
  • All 7 routing policies and exactly when to use each one
  • Health checks — endpoint monitoring, calculated health checks, and private resource patterns
  • How to use Route 53 with a domain registered on GoDaddy or another registrar
  • Hybrid DNS — Inbound and Outbound Resolver Endpoints for on-premises integration

Why this matters

At Zerodha, if the primary trading API region goes down at market open, Route 53 Failover routing detects the health check failure and automatically redirects traffic to the backup region — without any human intervention. At Swiggy, Latency-Based routing ensures users in Mumbai connect to ap-south-1 and users in Singapore connect to ap-southeast-1, getting the fastest response from the nearest region. At Hotstar, Weighted routing lets the team gradually shift 10% of traffic to a new version during a deployment and roll back instantly if error rates climb. These are not theoretical use cases — they are production patterns running at scale right now.

How DNS Resolution Works

Every time you type a domain name, a multi-step resolution process runs in milliseconds.

◈ DIAGRAM
You type: api.swiggy.com in your browser
Step 1: Browser checks local cache → not found
Step 2: Asks Local DNS Resolver (your ISP or 8.8.8.8)
Step 3: Local Resolver asks Root DNS Server
→ "I don't know api.swiggy.com, try .com TLD server"
Step 4: Local Resolver asks .com TLD Server
→ "I don't know, try swiggy.com Name Server"
Step 5: Local Resolver asks swiggy.com Name Server (authoritative)
→ "api.swiggy.com is at 13.235.45.67"
Step 6: Local Resolver returns IP to browser, caches it for the TTL
Step 7: Browser connects to 13.235.45.67

The whole process takes 20-100ms. Results are cached so subsequent requests skip most steps.

DNS terminology you must know:

Term What it means
Domain Registrar Where you buy domain names — Route 53, GoDaddy, Namecheap
DNS Records Instructions stored in DNS — A, AAAA, CNAME, NS, etc
Zone File The file containing all DNS records for a domain
Authoritative Name Server The server that has the definitive answer for your domain
TLD Top Level Domain — .com, .in, .org, .gov
SLD Second Level Domain — amazon.com, google.com
FQDN Fully Qualified Domain Name — api.www.example.com

What is Amazon Route 53

Route 53 is AWS's highly available, fully authoritative DNS service. Authoritative means you — the customer — can update DNS records directly. You are in full control of where traffic goes.

Route 53 is also a Domain Registrar. You can buy and register domain names directly through it, just like GoDaddy or Namecheap.

◈ DIAGRAM
Client types: myapp.in
→ Route 53 receives the DNS query
→ Looks up the record for myapp.in
→ Returns: 54.22.33.44
→ Client connects to EC2 at 54.22.33.44

Key facts:

TEXT
The only AWS service with a 100% availability SLA — AWS guarantees it never goes down
Can check health of your resources and route based on health
The name "Route 53" references port 53 — the standard DNS port
Cost: $0.50 per month per hosted zone

DNS Record Types

Each record is an instruction telling Route 53 how to respond to a query.

A Record — maps hostname to IPv4 address:

◈ DIAGRAM
example.com → 1.2.3.4

AAAA Record — maps hostname to IPv6 address:

◈ DIAGRAM
example.com → 2001:0db8:85a3::8a2e:0370:7334

CNAME Record — maps hostname to another hostname:

◈ DIAGRAM
www.example.com → app.example.com
Critical limitation: CANNOT create a CNAME for the root domain (Zone Apex)
Works: www.example.com → anything.com (subdomain — fine)
Fails: example.com → anything.com (root domain — not allowed)

NS Record — Name Server records:

Tell the internet which DNS servers are authoritative for your domain. When you buy a domain on GoDaddy and point it to Route 53, you update NS records at GoDaddy to use Route 53's name servers.

Must know deeply: A, AAAA, CNAME, NS

Hosted Zones

A Hosted Zone is a container in Route 53 holding all DNS records for a domain and its subdomains.

Cost: $0.50 per month per hosted zone

Public Hosted Zone — routes traffic on the public internet:

◈ DIAGRAM
Internet client → Public Hosted Zone (example.com) → EC2, ALB, CloudFront

Anyone on the internet can query a public hosted zone.

Private Hosted Zone — routes traffic only inside your VPC:

◈ DIAGRAM
EC2 inside VPC → Private Hosted Zone (company.internal)
db.company.internal → 10.0.0.35 (RDS instance private IP)
cache.company.internal → 10.0.10.9 (ElastiCache private IP)
api.company.internal → 10.0.5.20 (internal service)

Only resources inside your VPC can resolve these names. Microservices communicate using friendly names instead of hardcoded private IPs.

TTL — Time To Live

TTL is how long a DNS resolver caches a record before asking Route 53 again.

◈ DIAGRAM
High TTL (86400 seconds = 24 hours):
Fewer Route 53 queries — cheaper
Change a record → clients stuck with old value for up to 24 hours
Good for stable records that rarely change
Low TTL (60 seconds):
More Route 53 queries — slightly more expensive
Changes propagate within 60 seconds
Good when actively making changes or migrating

Safe strategy for changing a DNS record:

TEXT
Step 1: Lower TTL to 60 seconds
Step 2: Wait 24-48 hours (so all cached clients pick up the new low TTL)
Step 3: Make your DNS change
Step 4: Wait 60 seconds (all clients get the new value)
Step 5: Raise TTL back to 86400

If you skip Step 2, some clients still have the old high TTL cached and will not see your DNS change for hours.

Remember

TTL is mandatory on all records except Alias records. Alias records have no TTL — Route 53 manages it automatically.

CNAME vs Alias Records

AWS resources expose DNS hostnames, not static IPs:

TEXT
Your Load Balancer DNS: lb1-1234.ap-south-1.elb.amazonaws.com
You want users to access: api.myapp.in

You need to point your domain at this AWS hostname. Two options:

CNAME:

TEXT
Points a hostname to any other hostname
Only works for non-root subdomains — www.example.com works, example.com does not
Charges apply for DNS queries against CNAME records

Alias Record:

◈ DIAGRAM
Points a hostname specifically to an AWS resource
Works for BOTH root and non-root domains — example.com and www.example.com
Free — no query charges
Automatically tracks IP changes of the AWS resource (ALB IPs rotate — Alias handles transparently)
Built-in health check support
Always type A or AAAA — never CNAME type
No TTL — Route 53 manages it
CNAME:
www.example.com → lb.anything.com (works — non-root subdomain)
example.com → lb.anything.com (fails — root domain, Zone Apex)
Alias:
www.example.com → lb.ap-south-1.elb.amazonaws.com (works)
example.com → lb.ap-south-1.elb.amazonaws.com (works — root supported)

Valid Alias record targets:

TEXT
Elastic Load Balancers (ALB, NLB, CLB)
CloudFront Distributions
API Gateway
Elastic Beanstalk environments
S3 Websites
VPC Interface Endpoints
Global Accelerator
Route 53 record in the same hosted zone
Remember

You cannot set an Alias record pointing to an EC2 DNS name. EC2 is not a valid Alias target. Only the services listed above are supported.

Health Checks

Route 53 Health Checks monitor your resources and trigger automatic DNS failover when something goes wrong. Health checks work only for public resources — they cannot directly access private VPC resources.

Type 1 — Monitoring an Endpoint:

Route 53 sends health checker requests from 15 global locations to your endpoint.

TEXT
Protocol: HTTP, HTTPS, or TCP
Interval: 30 seconds (default) or 10 seconds (higher cost)
Healthy threshold: 3 consecutive successes
Healthy if: 18% or more of health checkers report healthy
Health check passes when: endpoint returns 2xx or 3xx status code
Can also check: specific text in the first 5,120 bytes of the response

Type 2 — Calculated Health Checks:

Combine multiple child health checks into one parent using AND/OR/NOT logic.

◈ DIAGRAM
Child health check A (region 1) ─┐
Child health check B (region 2) ─┼──> Parent health check (AND logic)
Child health check C (region 3) ─┘
Specify how many children must pass for parent to be healthy
Use: doing maintenance on one resource without failing overall health

Type 3 — Health Checks for Private Resources:

Route 53 health checkers run outside your VPC — they cannot reach private endpoints directly.

◈ DIAGRAM
Workaround:
EC2 in private subnet
↓
CloudWatch Metric (monitors CPU, errors, or custom metric)
↓
CloudWatch Alarm (fires when threshold crossed)
↓
Route 53 Health Check (monitors the CloudWatch Alarm state)

Route 53 indirectly monitors private resources through the CloudWatch Alarm state.

The 7 Routing Policies

1. Simple Routing — single resource, no health checks:

◈ DIAGRAM
Can return multiple IP values in the same record
If multiple values returned → client picks one at random
Cannot be associated with Health Checks
Client: "What is api.example.com?"
Route 53 returns: 11.22.33.44, 55.66.77.88, 99.11.22.33
Client picks one at random and connects

2. Weighted Routing — control traffic percentage:

◈ DIAGRAM
Assign relative weights to each record
Traffic % = record weight / sum of all weights
Can be associated with Health Checks
Weight = 0 → stop sending traffic to that resource
api.example.com → 11.22.33.44 (Weight: 70) → 70% of traffic
api.example.com → 55.66.77.88 (Weight: 20) → 20% of traffic
api.example.com → 99.11.22.33 (Weight: 10) → 10% of traffic
Use for: A/B testing new versions, gradual blue-green deployments,
distributing load across regions

3. Latency-Based Routing — route to lowest latency region:

TEXT
Latency measured between users and AWS Regions — not geographic distance
A user in Germany may be routed to us-east-1 if it has lower latency than eu-west-1
Can be associated with Health Checks
User in Mumbai:
Route 53 checks latency: ap-south-1=12ms, us-east-1=180ms, eu-west-1=210ms
Returns: ap-south-1 (lowest latency wins)

4. Failover Routing — active-passive disaster recovery:

◈ DIAGRAM
Health Check is mandatory on the primary record
If primary health check fails → Route 53 automatically returns secondary record
Normal: Client → Route 53 → Primary EC2 (healthy)
Failure: Primary health check fails
→ Route 53 automatically returns Secondary EC2
→ Client connects to Secondary (no manual action)

5. Geolocation Routing — route by physical location:

◈ DIAGRAM
Route based on where the user physically is — not latency
Specify routing by Continent, Country, or US State
Most specific location wins if rules overlap
Always create a Default record for users whose location matches nothing
User in Germany → EU server (German language content)
User in India → AP server (regional pricing and content)
User anywhere else → Default record (global fallback)
Use for: website localisation, content restriction by country, regional compliance
Remember

Geolocation routes by where the user physically is. Latency-based routes by which region responds fastest. A user physically close to a region may still have higher latency — they are different things.

6. Geoproximity Routing — shift traffic by adjusting bias:

◈ DIAGRAM
Route based on geographic location of users AND resources
Bias value shifts more or less traffic toward a specific resource
Positive bias (+1 to +99): expands coverage, attracts more traffic
Negative bias (-1 to -99): shrinks coverage, attracts less traffic
Requires Route 53 Traffic Flow feature
Scenario: gradually shift traffic from us-west-1 to us-east-1
Both at bias 0 → traffic split by proximity
us-east-1 bias = 50 → coverage expands → more traffic goes there
us-east-1 bias = 99 → almost all traffic shifted to us-east-1

7. IP-Based Routing — route by client IP address:

◈ DIAGRAM
You provide a list of CIDR ranges mapped to specific endpoints
Route 53 routes based on the client's IP address
User A (IP: 203.0.113.x) → matches CIDR 203.0.113.0/24 → EC2 in Region A
User B (IP: 200.5.4.x) → matches CIDR 200.5.4.0/24 → EC2 in Region B
Use for: route known ISP IP ranges to nearby endpoints,
reduce data transfer costs for specific networks

8. Multi-Value Routing — healthy IPs only:

◈ DIAGRAM
Returns up to 8 healthy records per query
Each record associated with a Health Check — only healthy included
Client receives multiple IPs and picks one
Not a substitute for a Load Balancer — client-side selection only
Client: "What is api.example.com?"
Route 53 checks health of all records:
192.0.2.2 → HEALTHY ← returned
198.51.100.2 → HEALTHY ← returned
203.0.113.2 → UNHEALTHY ← excluded
Returns only healthy IPs. Client picks one.

Using GoDaddy as Registrar with Route 53 for DNS

You can buy a domain on GoDaddy and manage DNS records on Route 53.

◈ DIAGRAM
1. Buy example.in on GoDaddy
2. Create a Public Hosted Zone in Route 53 for example.in
3. Route 53 gives you 4 Name Server (NS) records
4. Go to GoDaddy → Manage Domain → Change Nameservers
5. Paste Route 53's 4 NS records into GoDaddy
6. GoDaddy tells the internet: "Route 53 manages DNS for example.in"
7. All DNS queries for example.in now go to Route 53

From this point manage all DNS records inside Route 53 even though you bought the domain on GoDaddy.

Hybrid DNS — Resolver Endpoints

By default, EC2 instances inside your VPC can resolve AWS service DNS names and private hosted zone names. But when you have an on-premises data center connected via VPN or Direct Connect, DNS needs to work in both directions.

Inbound Endpoint — on-premises to AWS:

Bash
On-premises server needs to resolve: db.aws.internal
→ On-premises DNS Resolver sends query to Route 53 Inbound Endpoint
→ Route 53 looks up the Private Hosted Zone
→ Returns: db.aws.internal → 10.0.1.5
→ On-premises server connects to RDS in your VPC

Outbound Endpoint — AWS to on-premises:

◈ DIAGRAM
EC2 in VPC needs to resolve: erp.onpremise.internal
→ Route 53 Resolver: no match in any hosted zone
→ Forwards query to Outbound Endpoint
→ Outbound Endpoint queries on-premises DNS Resolvers via VPN/Direct Connect
→ On-premises DNS returns: erp.onpremise.internal → 192.168.1.10
→ EC2 connects to the on-premises ERP server

Both endpoints operate over VPN or Direct Connect to bridge DNS between AWS and on-premises.

Hands-on Lab — Hosted Zone, Records, and Weighted Routing

Step 1 — Create a Public Hosted Zone

◈ DIAGRAM
Route 53 → Hosted zones → Create hosted zone
Domain name: devops-lab.in (or any domain you own)
Type: Public hosted zone
Create hosted zone
Note the 4 NS records — update these at your domain registrar
to point DNS to Route 53.

Step 2 — Create a simple A record

◈ DIAGRAM
Route 53 → Hosted zones → devops-lab.in → Create record
Record name: www
Record type: A
Value: your EC2 public IP
TTL: 60 (low for testing — change to 300+ in production)
Create records
Test: dig www.devops-lab.in @8.8.8.8
Should return your EC2 IP within 60 seconds.

Step 3 — Create a Health Check

◈ DIAGRAM
Route 53 → Health checks → Create health check
Name: server-1-health
What to monitor: Endpoint
Protocol: HTTP Domain: your EC2 public IP Port: 80 Path: /
Request interval: 30 seconds Failure threshold: 3
Create health check
Wait 2 minutes → Health check status should show Healthy.

Step 4 — Create Weighted routing records

◈ DIAGRAM
Route 53 → Hosted zones → devops-lab.in → Create record
Record name: api
Record type: A
Routing policy: Weighted
Value: INSTANCE-1-IP Weight: 70 Record ID: server-1
Health check: server-1-health
Add another record:
Value: INSTANCE-2-IP Weight: 30 Record ID: server-2
Create records

Step 5 — Test traffic distribution

TEXT
Run multiple DNS lookups and observe which IP returns:
For 10 queries, approximately 7 should return Server 1 IP,
3 should return Server 2 IP.
dig api.devops-lab.in @8.8.8.8
(run several times — watch the answer change between the two IPs)

Step 6 — Test health check failover

◈ DIAGRAM
Stop the EC2 instance behind Server 1.
Wait 90 seconds for health check to detect failure (3 checks × 30s).
Route 53 → Health checks → server-1-health → Status: Unhealthy
Now all 10/10 DNS queries return Server 2 IP.
Route 53 automatically excluded the unhealthy endpoint.

Step 7 — Cleanup

◈ DIAGRAM
Route 53 → Hosted zones → devops-lab.in → select all records except NS and SOA
Delete records
Route 53 → Health checks → server-1-health → Delete
Route 53 → Hosted zones → devops-lab.in → Delete hosted zone

Production Best Practices and Common Pitfalls

  • Always use Alias records instead of CNAME for AWS resources — free, supports root domain, auto-tracks IP changes
  • Lower TTL before making DNS changes — if TTL is 24 hours and you change a record, clients are stuck for up to 24 hours
  • Always create a Default record with Geolocation routing — without it, users from unmatched locations get NXDOMAIN errors
  • Use health checks with Failover routing — Failover without health checks means Route 53 keeps routing to a dead endpoint
  • Use Latency-Based routing for multi-region deployments — it automatically finds the fastest region for each user
  • Do not confuse Geolocation with Latency-Based — Geolocation routes by where users are, Latency routes by where response is fastest

Quick Reference and Troubleshooting Commands

Task Command
List hosted zones aws route53 list-hosted-zones
Create hosted zone aws route53 create-hosted-zone --name example.com --caller-reference $(date +%s)
List records in zone aws route53 list-resource-record-sets --hosted-zone-id <ZONE-ID>
Create/update record aws route53 change-resource-record-sets --hosted-zone-id <ZONE-ID> --change-batch file://record.json
List health checks aws route53 list-health-checks
Get health check status aws route53 get-health-check-status --health-check-id <ID>
Delete health check aws route53 delete-health-check --health-check-id <ID>
Check DNS resolution dig api.example.com @8.8.8.8
Check which NS records serve a domain dig NS example.com

Common problems and fixes:

Problem Likely cause Fix
DNS change not propagating Old TTL still cached by resolvers Wait for old TTL to expire or lower TTL before future changes
NXDOMAIN for geolocation users No Default record created Add a Default record as fallback for unmatched locations
Failover not working Health check not associated with primary record Associate health check with primary record in the record set
Cannot create CNAME for root domain DNS standard limitation Use Alias record instead — supports root domain
Private hosted zone not resolving VPC not associated with the hosted zone Associate the VPC with the private hosted zone
Common Mistake

Creating a CNAME for the root domain (example.com). The DNS standard does not allow CNAME at the Zone Apex. Route 53 will reject it. Use an Alias record — it achieves the same result, works for root domains, and is free.

Common Mistake

Using Failover routing without health checks. Without a health check, Route 53 has no way to know the primary is down and will keep routing to a failed endpoint indefinitely. Health checks are mandatory for Failover routing to actually fail over.

Tip

For a production multi-region setup combine Latency-Based routing with health checks on every record. Latency-Based sends users to the fastest region normally. Health checks automatically exclude unhealthy regions. If ap-south-1 goes down during IPL streaming, Route 53 automatically routes users to the next fastest healthy region — no manual intervention, no ops team paged at 3 AM.

Common Mistakes to Avoid

Common Mistake

Using a CNAME for the root domain (example.com). DNS does not allow CNAME at the Zone Apex. Use an Alias record instead — it supports root domains, is free, and automatically tracks IP changes of the AWS resource.

Common Mistake

Using Failover routing without a health check on the primary record. Without a health check, Route 53 has no way to know the primary is down. It keeps routing to the failed endpoint indefinitely. Health checks are mandatory for Failover routing to actually fail over.

Tip

Lower your TTL to 60 seconds before making any DNS change. If TTL is 24 hours and you change a record, clients with the old value cached are stuck for up to 24 hours. Lower TTL → wait 24 hours for caches to expire → make the change → raise TTL back.

Resources

AWS Direct Connect vs Site-to-Site VPN Failover

AWS Direct Connect vs Site-to-Site VPN Failover

Direct Connect vs VPN isn't really either/or for production — it's a primary-plus-failover pattern. Here's how to design it, and when either/or is right.

5 min read•Aug 2026
Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot, at the crossover where Lambda stops being cheaper — 2026 pricing, invocation thresholds, and interruption math.

5 min read•Aug 2026
Secrets Manager vs Parameter Store vs Vault

Secrets Manager vs Parameter Store vs Vault

AWS Secrets Manager, Parameter Store, and HashiCorp Vault compared for 2026 - cost math, rotation, multi-cloud fit, and the Vault-to-OpenBao fork.

5 min read•Aug 2026
AWS VPC Security: Hardening Every Layer

AWS VPC Security: Hardening Every Layer

Most cloud security incidents start with a misconfigured VPC. Here's how to harden every layer — subnets, Security Groups, NACLs, and IAM — for production.

5 min read•Jul 2026
Event-Driven Architecture on AWS Explained

Event-Driven Architecture on AWS Explained

Event-driven architecture on AWS decouples services and absorbs traffic spikes using SQS, SNS, EventBridge, and Lambda — workflows that scale themselves.

5 min read•Jul 2026
S3 vs RDS vs DynamoDB: Choosing AWS Storage

S3 vs RDS vs DynamoDB: Choosing AWS Storage

Choosing S3, RDS, or DynamoDB wrong costs you in performance, cost, and scalability. Here is a practical decision guide based on your actual access patterns.

5 min read•Jul 2026
AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS bills surprise teams every month. Here are the 8 concrete actions that cut cloud spend by 40-60% without touching your application architecture.

5 min read•Jul 2026
EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2, Lambda, or Fargate — choosing the wrong AWS compute option costs you money and performance. Here is exactly when to use each one in production.

5 min read•Jul 2026

Explore More in AWS Networking and Security

All 6 Topics

Frequently Asked Questions

Is Amazon Route 53 - DNS Routing Policies for Production Architectures free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Amazon Route 53 - DNS Routing Policies for Production Architectures topic cover?

Configure Route 53 hosted zones, routing policies, health checks, and hybrid DNS endpoints to control global traffic routing and automatic failover.