What you will learn
- What an Auto Scaling Group is and the problem it solves
- The three numbers that every ASG is built around
- How Launch Templates tell ASG exactly what to launch
- How CloudWatch alarms trigger scaling decisions
- The four scaling policies and when to use each
- How ASG and Load Balancer work together for self-healing
- Cooldown periods and how to make them shorter
Why this matters
Imagine Swiggy running on 4 EC2 instances on a normal Tuesday. Then IPL final starts. 10x traffic hits in minutes. The 4 instances buckle, response times climb, users see errors. By the time an engineer wakes up and manually adds servers, the moment has passed.
Auto Scaling Groups solve this without any human involvement. The fleet grows automatically when traffic spikes and shrinks when traffic drops. And even on quiet days at 3 AM — if one instance crashes, ASG notices, terminates it, and launches a fresh replacement. Nobody wakes up.
That is the core idea. One setup. Permanent protection.
What is an Auto Scaling Group
An ASG is a managed fleet of EC2 instances. You define the rules. AWS handles the rest.
You say: "keep between 2 and 10 instances, target 50% CPU"ASG does: adds instances when CPU goes high removes instances when CPU drops replaces any instance that fails a health checkASG itself costs nothing. You only pay for the EC2 instances it creates.
The Three Numbers
Every ASG has exactly three numbers. Everything else flows from these.
Minimum : 2 Never go below this — ever Always at least 2 running for high availability If traffic drops to zero, still 2 instances running Desired : 4 Target count right now Scaling policies change this number up or down ASG adds or removes instances to match it Maximum : 10 Never go above this — ever Your cost protection ceiling Even during the biggest traffic spike, stops at 10When a scaling event fires, ASG changes the desired number first. Then it launches or terminates instances to match that number. Min and Max are hard walls desired can never cross.
RememberAlways set minimum to at least 2 and spread across 2 Availability Zones. Minimum of 1 means one AZ failure takes your whole app down until ASG can replace the instance — that takes minutes.
Launch Template — What Gets Launched
When ASG needs to add a new instance, it follows a Launch Template. You create this once and ASG uses it every time forever.
Launch Template saves: Which AMI to boot from Instance type (t3.medium, m5.large...) Security Groups IAM Role (what AWS services this instance can access) Key pair for SSH User Data script (runs once on first boot) Disk size and typeEvery new instance launched by ASG is identical. No manual setup. No snowflake servers.
RememberLaunch Configurations are the old version of Launch Templates and are deprecated. Always use Launch Templates — never Launch Configurations — for any new ASG.
Pre-baked AMI vs plain Amazon Linux
If your Launch Template uses a standard Amazon Linux AMI and installs your app via User Data, every new instance takes 5 to 10 minutes to become ready. During that time it cannot serve traffic.
Base AMI + User Data script: Instance boots → script runs apt install, npm install, builds app 5 to 10 minutes later → instance ready to serve traffic Custom AMI with app pre-installed: Instance boots → app already there, starts immediately 60 to 90 seconds later → instance ready to serve trafficFaster boot time means shorter cooldown. Shorter cooldown means ASG reacts faster to real spikes. Pre-baking your app into a custom AMI is one of the highest-impact changes you can make to an ASG setup.
ASG and Load Balancer Together
In production, ASG and Load Balancer are always used together. They do different jobs but depend on each other.
Load Balancer → distributes traffic, health checks every instanceASG → manages how many instances exist, replaces failuresThe teamwork looks like this:
ASG launches a new instance ↓Instance automatically registers with the Load Balancer Target Group ↓Load Balancer runs its health check ↓Instance passes → traffic starts flowing to itAnd when something goes wrong:
Instance starts failing health checks ↓Load Balancer stops sending traffic to it ↓ASG detects the unhealthy instance ↓ASG terminates it and launches a replacement ↓New instance passes health check → traffic flows againThis whole cycle takes 3 to 5 minutes. No engineer involved. This is self-healing infrastructure.
RememberUse health check type ELB, not EC2. EC2 health check only verifies the instance is running — not that your app is actually working. An instance can be up and returning errors on every request. ELB health check catches that. EC2 health check does not.
CloudWatch Alarms — How ASG Decides to Scale
ASG does not watch your traffic on its own. CloudWatch watches it and tells ASG what to do.
CloudWatch monitors: average CPU across all ASG instances Traffic spike → CPU climbs above 70% for 2 minutes ↓CloudWatch alarm fires ↓ASG scale-out policy triggers ↓Desired capacity changes: 4 → 6 ↓ASG launches 2 new instances ↓Load spreads, CPU drops back down Traffic quiet → CPU drops below 30% for 5 minutes ↓CloudWatch alarm fires ↓ASG scale-in policy triggers ↓Desired capacity changes: 6 → 4 ↓ASG terminates 2 instances cleanlyWhat to scale on:
| Metric | What it measures | Best for |
|---|---|---|
| CPUUtilization | Average CPU across all instances | Compute-heavy apps |
| RequestCountPerTarget | Requests per instance from ALB | Web APIs |
| ApproximateNumberOfMessages | SQS queue depth | Queue workers |
| Custom metric | Anything you push to CloudWatch | Business logic |
RequestCountPerTarget is excellent for APIs. Say "keep 1000 requests per instance" and ASG maintains that automatically by adding or removing instances. More meaningful than raw CPU for web workloads.
For queue-based workers, scale on ApproximateNumberOfMessages. Queue grows → scale out. Queue empties → scale in. CPU tells you nothing useful when your app is just reading from a queue.
The Four Scaling Policies
Target Tracking — start here for most things
You pick a target metric value. ASG automatically creates the CloudWatch alarms and figures out how many instances to add or remove. You do nothing else.
Example: keep average CPU at 50%ASG creates alarms, watches CPU, scales automaticallyTraffic drops CPU to 30% → ASG removes instancesTraffic pushes CPU to 70% → ASG adds instancesThis is the simplest policy and the right starting point for most workloads.
Step Scaling — different responses for different severity
A mild spike adds 2 instances. A severe spike adds 5. You define the thresholds and the steps.
CPU 50–70% → add 1 instance (mild)CPU 70–90% → add 3 instances (significant)CPU above 90% → add 5 instances (critical)Useful when a small traffic bump and a major traffic event need very different responses.
Scheduled Scaling — for predictable patterns
Zerodha knows market opens at 9:15 AM every weekday. Traffic spikes then. Every night it is quiet.
8:30 AM weekdays → set desired to 10, minimum to 84:00 PM weekdays → set desired to 3, minimum to 2Schedule it once. Never worry about it again.
Predictive Scaling — scale before traffic arrives
AWS looks at 14 days of your traffic history, predicts when the next surge will come, and pre-scales 30 minutes before it hits. Instead of reacting after the spike arrives and waiting 2-3 minutes for instances to boot, you are already ready.
Without predictive: spike arrives → alarm fires → instances boot → 2-3 min lagWith predictive: ASG pre-scales 30 min before → spike arrives → already readyWorks best when you have consistent daily or weekly patterns.
Cooldown Periods
After ASG launches new instances, it waits before launching more. This waiting time is the cooldown. Default is 300 seconds.
Why it exists:
Scale-out fires → 3 new instances start booting ↓Cooldown starts (300 seconds)CPU still looks high — new instances are still booting, not serving yetASG sees high CPU but does NOT act ↓Cooldown ends → instances running and serving → CPU normalisesASG checks again → no further action neededWithout cooldown, ASG would keep launching instances every 60 seconds while the first batch boots. You would end up with 3x more instances than needed.
TipPre-bake your app into a custom AMI so instances boot in 60-90 seconds instead of 8-10 minutes. Then set a 90-second cooldown instead of 300. ASG reacts faster, scales more precisely, and never over-provisions. This single change often cuts scaling overshoot by 60-70%.
Lifecycle Hooks — Pause Before Traffic
Sometimes a new instance needs to do something before it starts receiving traffic — register with a service discovery tool, pull secrets, warm up a cache.
Lifecycle hooks let you pause the instance at a specific point and run custom logic before it joins the fleet.
Instance boots ↓Lifecycle hook: instance pauses here (Pending:Wait state)Your code runs: pulls config from SSM, registers with Consul, warms cacheYou send "continue" signal ↓Instance moves to InServiceLoad Balancer starts sending trafficThe instance never receives a single request until your custom initialisation is complete.
Termination Policy — Which Instance Gets Removed
When ASG scales in, it needs to decide which instance to terminate. The default logic:
1. Balance across AZs first → remove from the AZ with the most instances2. Within that AZ → terminate the instance with the oldest Launch Template3. If tied → terminate the closest to its next billing hourYou can override with: OldestInstance, NewestInstance, OldestLaunchTemplate, or ClosestToNextInstanceHour.
Hands-on Lab — Create ASG, Test Self-Healing, Watch Scale-Out
This lab creates a working ASG, confirms self-healing works, and generates real CPU load to watch the scaling fire.
Step 1 — Create the Launch Template in the Console
EC2 → Launch Templates → Create Launch TemplateName: devops-web-ltAMI: Amazon Linux 2 (free tier eligible)Instance type: t2.microKey pair: your existing keySecurity Groups: your web security groupUser Data: paste the script belowyum update -yyum install -y nginx stressecho "<h1>Server: $(hostname -f)</h1>" > /usr/share/nginx/html/index.htmlsystemctl start nginxsystemctl enable nginxStep 2 — Create the Auto Scaling Group
EC2 → Auto Scaling Groups → Create Auto Scaling GroupName: devops-web-asgLaunch Template: devops-web-ltVPC: your VPCSubnets: select subnets in 2 different AZs Set group size: Minimum: 2 Desired: 2 Maximum: 5 Health check type: ELBHealth check grace period: 60 secondsStep 3 — Attach a Target Tracking Policy
In the ASG → Automatic Scaling → Add PolicyPolicy type: Target trackingMetric: Average CPU utilizationTarget value: 50%This creates CloudWatch alarms automatically. No extra work needed.
Step 4 — Verify both instances are healthy
EC2 → Auto Scaling Groups → devops-web-asg → Instance ManagementBoth instances should show: Lifecycle = InService, Health = HealthyStep 5 — Test self-healing
EC2 → Instances → select one of the ASG instances → Stop Instance Wait 60 seconds then check:EC2 → Auto Scaling Groups → devops-web-asg → Activity You will see: "Terminating EC2 instance: i-0abc123..." "Launching a new EC2 instance..." New instance registers and passes health checkStep 6 — Generate CPU load and watch scale-out
SSH into one of the running instances:ssh -i your-key.pem ec2-user@INSTANCE-IP Run the stress tool for 5 minutes:stress --cpu 4 --timeout 300 In the console, watch:EC2 → Auto Scaling Groups → devops-web-asg → Monitoring → CPU Utilization After 2-3 minutes CPU exceeds 50%. The Target Tracking alarm fires.ASG launches more instances. Watch Activity History update in real time. When stress ends, CPU drops. After the cooldown period ASG scales back in.Step 7 — Cleanup
EC2 → Auto Scaling Groups → devops-web-asg → DeleteCheck "Force delete" to terminate all instances immediatelyEC2 → Launch Templates → devops-web-lt → DeleteCommon Mistakes to Avoid
Common MistakeSetting minimum to 1. One instance means one Availability Zone. If that AZ has hardware trouble your app goes down until ASG replaces the instance — which takes minutes. Always set minimum to at least 2 across 2 AZs. High availability requires at least 2 copies.
Common MistakeUsing EC2 health check type instead of ELB. An instance can be running perfectly but your application crashing on every request. EC2 health check calls it healthy. ELB health check calls it unhealthy and ASG replaces it. Always use ELB health check for application workloads.
Common MistakeUsing a plain Amazon Linux AMI when you could pre-bake. Every 8-minute boot time is 8 minutes your new instances cannot serve traffic. Pre-bake your app. Boot in 90 seconds. Cooldown drops from 300 to 90 seconds. ASG reacts 3x faster to real spikes.
SecuritySet Maximum to a number your budget can handle. During a traffic spike or a bug causing infinite scaling, the Maximum capacity is the only thing protecting your bill. Calculate your maximum based on cost, not just technical need.
TipFor queue-based workers reading from SQS — scale on
ApproximateNumberOfMessagesnot CPU. Set a target of messages-per-instance (say 1000 per worker). Queue builds up → instances scale out. Queue drains → instances scale in. Perfect matching between work and compute. CPU is irrelevant for a queue consumer.
Instance Refresh — Rolling Updates Without Downtime
When you update your Launch Template — new AMI, new instance type, new User Data — existing instances are still running the old version. Instance Refresh replaces them gradually without taking your application offline.
You update Launch Template to v2 (new AMI with latest app) ↓Trigger Instance Refresh with 90% minimum healthy threshold ↓ASG replaces one instance at a timeWaits for new instance to pass health check before touching the next ↓Eventually all instances running new Launch Template version ↓Zero downtime. No manual work. EC2 → Auto Scaling Groups → devops-web-asg → Instance Refresh → StartTipAlways trigger an Instance Refresh after updating a Launch Template. Without it, old instances keep running the old version indefinitely. Only new scale-out instances use the new template.