What you will learn
- Why load balancers exist — the four problems they solve that a single EC2 cannot
- ALB — Layer 7 smart routing by path, hostname, query string, and headers
- How ALB dynamic port mapping solves the container port conflict problem
- NLB — Layer 4 extreme performance, static IP, and when it beats ALB
- GWLB — Layer 3 for transparent security appliance inspection
- Health checks — how to design a /health endpoint that actually means healthy
- The correct Security Group setup so EC2 is never directly reachable from the internet
- Sticky sessions — the three cookie types and why ElastiCache is the better solution
- Cross-zone load balancing — defaults and costs per LB type
- SSL termination and SNI for multiple domains on one ALB
- Connection Draining — preventing dropped requests during deployments
- Auto Scaling Group integration — how instances register and deregister automatically
Why this matters
During a Swiggy Big Billion Day sale, traffic spikes 10x in under a minute. A single EC2 instance cannot handle this — it buckles and crashes. A load balancer distributes traffic across many EC2 instances, health-checks each one, stops routing to any that fail, and integrates with Auto Scaling Groups to add instances automatically as load grows. The load balancer is the entry point to every production application. Understanding which type to use, how to configure health checks correctly, and how SSL termination works is fundamental to building any AWS architecture.
Why Load Balancers Exist
Your application runs on a single EC2 instance. Traffic spikes. The instance is maxed out. Response times climb. It crashes. Every user sees a 502 error.
Even on a normal day — what happens when that EC2 needs a restart for a security patch? Every user gets an outage.
A load balancer sits in front of your application servers and forwards incoming requests to them. Users never talk to EC2 directly. They talk to the load balancer, and the load balancer decides which backend handles each request.
What a load balancer gives you:
Traffic distribution → no single server overwhelmedSingle DNS entry → one address like api.swiggy.in regardless of how many servers run behindHealth checking → LB checks each server continuously, stops routing to failed ones automaticallySSL termination → HTTPS handled at LB, EC2 only deals with plain HTTP internallyHigh availability → runs across multiple AZs, traffic routes to healthy AZs if one failsSeparation of public and private → EC2 stays private, only the LB is public-facingWhy use AWS ELB instead of running your own nginx or HAProxy:
You could self-manage a load balancer on EC2. The problem: you now own that server — patches, scaling, certificates, availability. AWS ELB is fully managed. AWS guarantees availability, handles upgrades, and integrates natively with EC2, ASG, ECS, ACM, CloudWatch, Route 53, WAF, and Global Accelerator.
AWS has 4 load balancer types:
ALB — Application Load Balancer (2016) Layer 7 — HTTP/HTTPS Smart routing. Best for web apps, microservices, containers. NLB — Network Load Balancer (2017) Layer 4 — TCP/UDP Extreme performance. Static IP. Non-HTTP traffic. GWLB — Gateway Load Balancer (2020) Layer 3 — IP Packets Security appliances. Firewalls, IDS/IPS inspection. CLB — Classic Load Balancer (2009) Layer 4 + 7 Deprecated. Do not use for new workloads.Health Checks and Security Groups
How the LB knows who is alive:
The LB checks every backend instance on a schedule using a protocol, port, and path you configure.
LB checks every instance every 30 seconds: Protocol: HTTP Port: 80 Path: /health 200 OK → instance is healthy → keep sending traffic Non-200 → instance is unhealthy → stop sending after 3 failures Recovers → 200 OK again → instance added back automaticallyDesign a meaningful /health endpoint:
## Good /health endpoint — checks actual dependenciesdef health(): checks = {} ## Can we connect to the database? try: db.execute("SELECT 1") checks['database'] = 'ok' except Exception: checks['database'] = 'failed' return jsonify(checks), 503 ## Can we connect to Redis? try: redis.ping() checks['redis'] = 'ok' except Exception: checks['redis'] = 'failed' return jsonify(checks), 503 return jsonify(checks), 200 ## Returns 200 only when ALL dependencies are healthy ## LB sees non-200 → removes instance from rotation automaticallyThe correct Security Group setup:
Internet ↓ port 80/443 from 0.0.0.0/0Load Balancer Security Group ↓ port 80 from Load Balancer Security Group ID ONLYEC2 Security Group → EC2 Instance Internet → EC2 directly → BLOCKEDInternet → LB → EC2 → ALLOWED| Security Group | Rule | Source |
|---|---|---|
| LB SG | Allow 80, 443 | 0.0.0.0/0 (anywhere) |
| EC2 SG | Allow 80 | LB Security Group ID only |
SecurityEC2 SG allowing only from LB SG is the most secure setup. Never allow direct internet access to EC2 when a load balancer is in front. If you allow 0.0.0.0/0 on EC2, the load balancer provides no security benefit — anyone can bypass it and hit EC2 directly.
Application Load Balancer — Layer 7
ALB operates at Layer 7. It reads and understands your HTTP request — the URL path, hostname, query string, and headers — and makes intelligent routing decisions.
One ALB, multiple services: myapp.in/users → Target Group: User Service (3 EC2 instances) myapp.in/orders → Target Group: Order Service (2 EC2 instances) myapp.in/search → Target Group: Search Service (1 Lambda function) Without ALB: 3 separate load balancers = 3x the costWith ALB: 1 LB, path-based routing, fraction of the costALB supports HTTP/2, WebSocket, and HTTP-to-HTTPS redirects natively.
4 routing rules:
Path-based routing:
myapp.in/users → User Servicemyapp.in/orders → Order Servicemyapp.in/admin → Admin Servicemyapp.in/* → Default catch-allHostname-based routing:
api.example.in → API Target Groupadmin.example.in → Admin Target Groupstatic.example.in → S3 or CDN Target GroupQuery string routing:
example.in/app?Platform=Mobile → Mobile Backend (EC2 on AWS)example.in/app?Platform=Desktop → Desktop Backend (On-premises via VPN)HTTP header routing:
User-Agent: MobileApp → Mobile backendUser-Agent: Browser → Web backendX-Internal: true → Internal service backendALB Target Types:
| Target Type | Description |
|---|---|
| EC2 Instances | Standard servers, managed by ASG |
| ECS Tasks | Containers with dynamic port mapping |
| Lambda Functions | ALB converts HTTP request to JSON event |
| IP Addresses | Private IPs only — on-premises servers via VPN |
Dynamic port mapping for containers:
When multiple containers run on the same EC2 host, each gets a random port at launch. ALB handles this automatically — when an ECS task starts, it registers with the ALB Target Group along with whatever port it was assigned. ALB updates routing. When the container stops, ALB removes it. You never touch port configuration manually.
One EC2 host running 3 containers: Container A → assigned port 32768 → registered with ALB TG Container B → assigned port 32769 → registered with ALB TG Container C → assigned port 32770 → registered with ALB TG ALB routes to each on its exact port automaticallyClient IP — the X-Forwarded-For header:
When ALB forwards a request to EC2, the instance sees the ALB's IP, not the real user IP. The actual client IP is in the X-Forwarded-For header.
Client IP: 12.34.56.78 → ALB adds: X-Forwarded-For: 12.34.56.78 (original client IP) X-Forwarded-Port: 443 (original port) X-Forwarded-Proto: https (original protocol) → EC2 reads X-Forwarded-For to get the real client IPRememberIf your application logs IP addresses for analytics or rate limiting, it must read X-Forwarded-For, not the connection source IP. The source IP is always the ALB's private IP. NLB does not have this problem — it passes the real client IP through directly.
NLB and GWLB — Layer 4 and Layer 3
Network Load Balancer — Layer 4:
A gaming company runs servers communicating over TCP. They need 5 million concurrent connections with sub-millisecond latency. Their enterprise clients must whitelist specific IP addresses in their firewalls. ALB cannot solve either of these problems.
NLB operates at Layer 4. It forwards TCP and UDP packets directly to targets without reading HTTP content.
Use NLB when: Need a static predictable IP (client firewall whitelisting) Traffic is TCP or UDP, not HTTP Extreme throughput required (millions req/sec) Ultra-low latency critical (gaming, financial trading, IoT) Need NLB in front of ALB (static IP + smart HTTP routing combined) Use ALB instead when: Web apps with HTTP/HTTPS Path or hostname-based routing needed Microservices or container workloadsKey NLB facts:
One static IP per Availability Zone — can assign Elastic IPsApproximately 100 microseconds latency (ALB is ~1ms or more)Millions of requests per secondDoes NOT add X-Forwarded-For header — backend sees real client IP directlyDoes NOT use Security Groups by default — access control on EC2 Security GroupsHealth checks support TCP, HTTP, and HTTPS protocolsNLB Target Groups:
| Target Type | Notes |
|---|---|
| EC2 Instances | Standard servers |
| IP Addresses | Private IPs only — on-premises via VPN |
| ALB | Chain NLB → ALB to get static IP + Layer 7 routing |
Gateway Load Balancer — Layer 3:
Your company processes financial transactions. Compliance requires all traffic to pass through a third-party network firewall before reaching your application.
GWLB sits at Layer 3. Every packet entering your network routes through GWLB, which sends it to a fleet of security appliances. If the packet passes inspection, GWLB forwards it. If malicious, the appliance drops it.
Traffic flow with GWLB: Internet users → Route Table (routes all traffic to GWLB) → Gateway Load Balancer → Target Group: Security Appliances (3rd party firewall fleet) → Safe traffic → forwarded to your application → Malicious traffic → dropped by applianceGWLB uses the GENEVE protocol on port 6081 to encapsulate traffic between itself and the appliance fleet. Target groups contain EC2 instances running security software.
LB type comparison:
| Feature | ALB | NLB | GWLB | CLB |
|---|---|---|---|---|
| Layer | 7 (HTTP) | 4 (TCP/UDP) | 3 (IP) | 4 + 7 |
| Protocol | HTTP, HTTPS, gRPC | TCP, UDP, TLS | IP packets (GENEVE) | TCP, HTTP |
| Static IP | No | Yes (per AZ) | No | No |
| Smart routing | Yes (path, host, query) | No | No | No |
| Use case | Web apps, microservices | Gaming, IoT, trading | Firewalls, IDS/IPS | Legacy only |
Advanced Features
Sticky Sessions (Session Affinity):
A user adds items to their cart. Cart is stored in EC2 memory, not a database. The next request goes to a different instance — the cart is gone. Stickiness ensures the same user always routes to the same EC2 instance.
Three cookie types for ALB:
Custom cookie → generated by your application, can include custom attributesApplication cookie → generated by ALB itself, name: AWSALBAPPDuration-based → generated by LB, name: AWSALB, you set expirationCommon MistakeStickiness causes uneven load distribution. If most sticky users end up on one instance, you lose the benefit of load balancing. The better architecture is storing session data externally in ElastiCache Redis — so any instance can serve any user and stickiness is not needed at all.
Cross-Zone Load Balancing:
Without cross-zone, each LB node distributes traffic only to instances in its own AZ. With cross-zone, each LB node distributes evenly across all instances in all AZs.
| LB Type | Cross-Zone Default | Inter-AZ Data Cost |
|---|---|---|
| ALB | Always ON | No charge |
| NLB | OFF by default | Charged if enabled |
| GWLB | OFF by default | Charged if enabled |
RememberALB cross-zone is always on and free. NLB and GWLB charge for inter-AZ data transfer when cross-zone is enabled.
SSL Termination and SNI:
User → HTTPS (encrypted, public internet) → Load Balancer (SSL termination here) → HTTP (unencrypted, private VPC) → EC2 InstanceThe LB uses an X.509 certificate from ACM. EC2 instances receive plain HTTP — no SSL cert management on servers.
SNI (Server Name Indication) lets one ALB serve multiple HTTPS domains. The client tells the server which domain it is connecting to at the start of the TLS handshake. ALB reads the SNI hostname and picks the correct certificate.
shop.example.in → shop certificateapi.example.in → api certificateadmin.example.in → admin certificateAll on ONE ALB, one IP — SNI picks the right cert automatically| LB Type | Multiple SSL Certs |
|---|---|
| CLB (v1) | One SSL cert only — multiple domains = multiple CLBs |
| ALB (v2) | Multiple SSL certs via SNI — multiple domains = one ALB |
| NLB (v2) | Multiple SSL certs via SNI — multiple domains = one NLB |
Connection Draining (Deregistration Delay):
During a deployment you deregister an EC2 instance. That instance is currently handling 200 in-flight requests — users mid-checkout, uploads in progress. Terminating immediately drops all 200.
Connection Draining tells the LB to stop sending new requests to the instance being removed while allowing existing requests to complete.
You deregister Instance B: New requests → Instance A and C only (B excluded immediately) Existing requests on Instance B: → still being processed → LB waits up to 300 seconds (default) → requests complete → instance safely terminatedConfigurable from 1 to 3600 seconds. Default is 300 seconds.
Low (10-30 seconds) → short-lived REST APIsHigh (300+ seconds) → long operations like file uploads or video processingSet to 0 → disable — instance removed immediately, in-flight requests failRememberThe feature is called "Connection Draining" for CLB and "Deregistration Delay" for ALB and NLB. Same behaviour, different name — a very common confusion point.
Hands-on Lab — ALB with Two Target Instances
Step 1 — Launch two EC2 instances
EC2 → Launch Instances → Launch instance (do this twice) Instance 1:Name: server-1a AMI: Amazon Linux 2023 Type: t2.microSubnet: public subnet ap-south-1aUser Data:yum install -y nginxecho "<h1>Server 1 — ap-south-1a</h1>" > /usr/share/nginx/html/index.htmlsystemctl start nginxInstance 2: same but subnet ap-south-1b, User Data says "Server 2 — ap-south-1b"Step 2 — Create Security Groups
EC2 → Security Groups → Create security groupName: alb-sg Inbound: HTTP 80 from 0.0.0.0/0 Create Create security groupName: ec2-sg Inbound: HTTP 80 from alb-sg (select the SG, not an IP)This means only the ALB can reach EC2 — internet cannot bypass the LB. Assign ec2-sg to both EC2 instances:EC2 → select instance → Actions → Security → Change security groupsStep 3 — Create Target Group
EC2 → Target Groups → Create target groupType: Instances Name: devops-tg Protocol: HTTP Port: 80Health check path: / Healthy threshold: 2Next → select both instances → Include as pending → Create target groupStep 4 — Create the ALB
EC2 → Load Balancers → Create load balancer → Application Load BalancerName: devops-alb Scheme: Internet-facingSubnets: select public-1a and public-1bSecurity group: alb-sgListener: HTTP 80 → Forward to devops-tgCreate load balancer — takes 2-3 minutesStep 5 — Test distribution
EC2 → Load Balancers → devops-alb → DNS name (copy it)Open browser → http://PASTE-ALB-DNS-HERERefresh several times — response alternates between Server 1 and Server 2.Step 6 — Test health check failover
Stop server-1a instance → wait 30 seconds for ALB to detect failureRefresh browser multiple times → all responses show Server 2 onlyStart server-1a → wait 30 seconds → Server 1 rejoins rotationStep 7 — Cleanup
EC2 → Load Balancers → devops-alb → Actions → DeleteEC2 → Target Groups → devops-tg → Actions → DeleteEC2 → Instances → terminate both instancesEC2 → Security Groups → delete alb-sg and ec2-sgProduction Best Practices and Common Pitfalls
- Always restrict EC2 Security Groups to allow only from the LB Security Group — never 0.0.0.0/0 on EC2
- Design /health endpoints to check actual dependencies — a 200 response from an app that cannot reach the database is worse than useless
- Use ALB for all HTTP/HTTPS workloads — do not use NLB unless you specifically need static IPs or extreme throughput
- Use HTTPS listeners with ACM certificates on all internet-facing ALBs — never serve plain HTTP in production
- Add a redirect rule on port 80 to forward all HTTP to HTTPS instead of having a plain HTTP listener
- Use ElastiCache for session storage instead of sticky sessions — sticky sessions cause uneven load distribution and are a scaling bottleneck
- Set Deregistration Delay based on your actual request duration — too high leaves instances draining unnecessarily, too low drops in-flight requests
Quick Reference and Troubleshooting Commands
| Task | Command |
|---|---|
| List all load balancers | aws elbv2 describe-load-balancers --region ap-south-1 |
| List target groups | aws elbv2 describe-target-groups --region ap-south-1 |
| Check target health | aws elbv2 describe-target-health --target-group-arn <ARN> --region ap-south-1 |
| Register instance with TG | aws elbv2 register-targets --target-group-arn <ARN> --targets Id=<i-xxx> |
| Deregister instance from TG | aws elbv2 deregister-targets --target-group-arn <ARN> --targets Id=<i-xxx> |
| Describe listeners | aws elbv2 describe-listeners --load-balancer-arn <ARN> --region ap-south-1 |
| Describe listener rules | aws elbv2 describe-rules --listener-arn <ARN> --region ap-south-1 |
Common problems and fixes:
| Problem | Likely cause | Fix |
|---|---|---|
| All targets unhealthy | Health check path returning non-200 or port blocked | Verify /health endpoint works, check EC2 SG allows port from ALB SG |
| 504 Gateway Timeout | Target not responding within timeout | Increase timeout in target group settings, check app performance |
| Traffic not distributing evenly | Sticky sessions enabled or cross-zone disabled | Check sticky session config, enable cross-zone load balancing |
| EC2 directly reachable from internet | EC2 SG allowing 0.0.0.0/0 instead of ALB SG | Update EC2 Security Group source to use ALB Security Group ID |
| SSL cert not found on Edge-Optimized API | Cert in wrong region | Edge-Optimized requires cert in us-east-1 specifically |
| Connection drops during deployment | Deregistration Delay too short | Increase to match longest request duration |
Common MistakeLeaving port 22 or port 80 open to 0.0.0.0/0 on EC2 Security Groups when a load balancer is in front. This defeats the entire purpose of the load balancer from a security perspective — attackers can bypass it and hit EC2 directly. EC2 Security Groups should only allow traffic from the LB Security Group ID.
TipWhen debugging why a target is unhealthy, always start with the target health endpoint in the ALB console — it shows the exact reason for the last failed health check including the HTTP status code and response body. This is far faster than SSH-ing into the instance and checking logs manually.
Common Mistakes to Avoid
Common MistakeLeaving port 80 or 443 open to 0.0.0.0/0 on EC2 Security Groups when an ALB is in front. This defeats the load balancer from a security perspective — attackers bypass the ALB and hit EC2 directly. EC2 SG should allow traffic only from the ALB Security Group ID.
Common MistakeDesigning a /health endpoint that always returns 200 regardless of actual state. If your DB is down but /health says 200, the LB keeps routing traffic to a broken instance. Design /health to check real dependencies.
TipUse ElastiCache for session storage instead of sticky sessions. Sticky sessions cause uneven load distribution — the majority of users end up routed to the same instance while others sit idle. Store sessions in Redis and any instance can serve any user.