Skip to main content

Auto Scaling Groups - Dynamic Scaling for Production Workloads

Configure Auto Scaling Groups with Launch Templates, scaling policies, and CloudWatch alarms to automatically adjust EC2 fleet size based on real demand.

What you will learn

  • What an Auto Scaling Group is and the problem it solves
  • The three numbers that every ASG is built around
  • How Launch Templates tell ASG exactly what to launch
  • How CloudWatch alarms trigger scaling decisions
  • The four scaling policies and when to use each
  • How ASG and Load Balancer work together for self-healing
  • Cooldown periods and how to make them shorter

Why this matters

Imagine Swiggy running on 4 EC2 instances on a normal Tuesday. Then IPL final starts. 10x traffic hits in minutes. The 4 instances buckle, response times climb, users see errors. By the time an engineer wakes up and manually adds servers, the moment has passed.

Auto Scaling Groups solve this without any human involvement. The fleet grows automatically when traffic spikes and shrinks when traffic drops. And even on quiet days at 3 AM — if one instance crashes, ASG notices, terminates it, and launches a fresh replacement. Nobody wakes up.

That is the core idea. One setup. Permanent protection.

What is an Auto Scaling Group

An ASG is a managed fleet of EC2 instances. You define the rules. AWS handles the rest.

TEXT
You say: "keep between 2 and 10 instances, target 50% CPU"
ASG does: adds instances when CPU goes high
removes instances when CPU drops
replaces any instance that fails a health check

ASG itself costs nothing. You only pay for the EC2 instances it creates.

The Three Numbers

Every ASG has exactly three numbers. Everything else flows from these.

TEXT
Minimum : 2
Never go below this — ever
Always at least 2 running for high availability
If traffic drops to zero, still 2 instances running
Desired : 4
Target count right now
Scaling policies change this number up or down
ASG adds or removes instances to match it
Maximum : 10
Never go above this — ever
Your cost protection ceiling
Even during the biggest traffic spike, stops at 10

When a scaling event fires, ASG changes the desired number first. Then it launches or terminates instances to match that number. Min and Max are hard walls desired can never cross.

Remember

Always set minimum to at least 2 and spread across 2 Availability Zones. Minimum of 1 means one AZ failure takes your whole app down until ASG can replace the instance — that takes minutes.

Launch Template — What Gets Launched

When ASG needs to add a new instance, it follows a Launch Template. You create this once and ASG uses it every time forever.

TEXT
Launch Template saves:
Which AMI to boot from
Instance type (t3.medium, m5.large...)
Security Groups
IAM Role (what AWS services this instance can access)
Key pair for SSH
User Data script (runs once on first boot)
Disk size and type

Every new instance launched by ASG is identical. No manual setup. No snowflake servers.

Remember

Launch Configurations are the old version of Launch Templates and are deprecated. Always use Launch Templates — never Launch Configurations — for any new ASG.

Pre-baked AMI vs plain Amazon Linux

If your Launch Template uses a standard Amazon Linux AMI and installs your app via User Data, every new instance takes 5 to 10 minutes to become ready. During that time it cannot serve traffic.

Bash
Base AMI + User Data script:
Instance boots → script runs apt install, npm install, builds app
5 to 10 minutes later → instance ready to serve traffic
Custom AMI with app pre-installed:
Instance boots → app already there, starts immediately
60 to 90 seconds later → instance ready to serve traffic

Faster boot time means shorter cooldown. Shorter cooldown means ASG reacts faster to real spikes. Pre-baking your app into a custom AMI is one of the highest-impact changes you can make to an ASG setup.

ASG and Load Balancer Together

In production, ASG and Load Balancer are always used together. They do different jobs but depend on each other.

◈ DIAGRAM
Load Balancer → distributes traffic, health checks every instance
ASG → manages how many instances exist, replaces failures

The teamwork looks like this:

◈ DIAGRAM
ASG launches a new instance
↓
Instance automatically registers with the Load Balancer Target Group
↓
Load Balancer runs its health check
↓
Instance passes → traffic starts flowing to it

And when something goes wrong:

◈ DIAGRAM
Instance starts failing health checks
↓
Load Balancer stops sending traffic to it
↓
ASG detects the unhealthy instance
↓
ASG terminates it and launches a replacement
↓
New instance passes health check → traffic flows again

This whole cycle takes 3 to 5 minutes. No engineer involved. This is self-healing infrastructure.

Remember

Use health check type ELB, not EC2. EC2 health check only verifies the instance is running — not that your app is actually working. An instance can be up and returning errors on every request. ELB health check catches that. EC2 health check does not.

CloudWatch Alarms — How ASG Decides to Scale

ASG does not watch your traffic on its own. CloudWatch watches it and tells ASG what to do.

◈ DIAGRAM
CloudWatch monitors: average CPU across all ASG instances
Traffic spike → CPU climbs above 70% for 2 minutes
↓
CloudWatch alarm fires
↓
ASG scale-out policy triggers
↓
Desired capacity changes: 4 → 6
↓
ASG launches 2 new instances
↓
Load spreads, CPU drops back down
Traffic quiet → CPU drops below 30% for 5 minutes
↓
CloudWatch alarm fires
↓
ASG scale-in policy triggers
↓
Desired capacity changes: 6 → 4
↓
ASG terminates 2 instances cleanly

What to scale on:

Metric What it measures Best for
CPUUtilization Average CPU across all instances Compute-heavy apps
RequestCountPerTarget Requests per instance from ALB Web APIs
ApproximateNumberOfMessages SQS queue depth Queue workers
Custom metric Anything you push to CloudWatch Business logic

RequestCountPerTarget is excellent for APIs. Say "keep 1000 requests per instance" and ASG maintains that automatically by adding or removing instances. More meaningful than raw CPU for web workloads.

For queue-based workers, scale on ApproximateNumberOfMessages. Queue grows → scale out. Queue empties → scale in. CPU tells you nothing useful when your app is just reading from a queue.

The Four Scaling Policies

Target Tracking — start here for most things

You pick a target metric value. ASG automatically creates the CloudWatch alarms and figures out how many instances to add or remove. You do nothing else.

◈ DIAGRAM
Example: keep average CPU at 50%
ASG creates alarms, watches CPU, scales automatically
Traffic drops CPU to 30% → ASG removes instances
Traffic pushes CPU to 70% → ASG adds instances

This is the simplest policy and the right starting point for most workloads.

Step Scaling — different responses for different severity

A mild spike adds 2 instances. A severe spike adds 5. You define the thresholds and the steps.

◈ DIAGRAM
CPU 50–70% → add 1 instance (mild)
CPU 70–90% → add 3 instances (significant)
CPU above 90% → add 5 instances (critical)

Useful when a small traffic bump and a major traffic event need very different responses.

Scheduled Scaling — for predictable patterns

Zerodha knows market opens at 9:15 AM every weekday. Traffic spikes then. Every night it is quiet.

◈ DIAGRAM
8:30 AM weekdays → set desired to 10, minimum to 8
4:00 PM weekdays → set desired to 3, minimum to 2

Schedule it once. Never worry about it again.

Predictive Scaling — scale before traffic arrives

AWS looks at 14 days of your traffic history, predicts when the next surge will come, and pre-scales 30 minutes before it hits. Instead of reacting after the spike arrives and waiting 2-3 minutes for instances to boot, you are already ready.

◈ DIAGRAM
Without predictive: spike arrives → alarm fires → instances boot → 2-3 min lag
With predictive: ASG pre-scales 30 min before → spike arrives → already ready

Works best when you have consistent daily or weekly patterns.

Cooldown Periods

After ASG launches new instances, it waits before launching more. This waiting time is the cooldown. Default is 300 seconds.

Why it exists:

◈ DIAGRAM
Scale-out fires → 3 new instances start booting
↓
Cooldown starts (300 seconds)
CPU still looks high — new instances are still booting, not serving yet
ASG sees high CPU but does NOT act
↓
Cooldown ends → instances running and serving → CPU normalises
ASG checks again → no further action needed

Without cooldown, ASG would keep launching instances every 60 seconds while the first batch boots. You would end up with 3x more instances than needed.

Tip

Pre-bake your app into a custom AMI so instances boot in 60-90 seconds instead of 8-10 minutes. Then set a 90-second cooldown instead of 300. ASG reacts faster, scales more precisely, and never over-provisions. This single change often cuts scaling overshoot by 60-70%.

Lifecycle Hooks — Pause Before Traffic

Sometimes a new instance needs to do something before it starts receiving traffic — register with a service discovery tool, pull secrets, warm up a cache.

Lifecycle hooks let you pause the instance at a specific point and run custom logic before it joins the fleet.

◈ DIAGRAM
Instance boots
↓
Lifecycle hook: instance pauses here (Pending:Wait state)
Your code runs: pulls config from SSM, registers with Consul, warms cache
You send "continue" signal
↓
Instance moves to InService
Load Balancer starts sending traffic

The instance never receives a single request until your custom initialisation is complete.

Termination Policy — Which Instance Gets Removed

When ASG scales in, it needs to decide which instance to terminate. The default logic:

◈ DIAGRAM
1. Balance across AZs first → remove from the AZ with the most instances
2. Within that AZ → terminate the instance with the oldest Launch Template
3. If tied → terminate the closest to its next billing hour

You can override with: OldestInstance, NewestInstance, OldestLaunchTemplate, or ClosestToNextInstanceHour.

Hands-on Lab — Create ASG, Test Self-Healing, Watch Scale-Out

This lab creates a working ASG, confirms self-healing works, and generates real CPU load to watch the scaling fire.

Step 1 — Create the Launch Template in the Console

◈ DIAGRAM
EC2 → Launch Templates → Create Launch Template
Name: devops-web-lt
AMI: Amazon Linux 2 (free tier eligible)
Instance type: t2.micro
Key pair: your existing key
Security Groups: your web security group
User Data: paste the script below
Bash
#!/bin/bash
yum update -y
yum install -y nginx stress
echo "<h1>Server: $(hostname -f)</h1>" > /usr/share/nginx/html/index.html
systemctl start nginx
systemctl enable nginx

Step 2 — Create the Auto Scaling Group

◈ DIAGRAM
EC2 → Auto Scaling Groups → Create Auto Scaling Group
Name: devops-web-asg
Launch Template: devops-web-lt
VPC: your VPC
Subnets: select subnets in 2 different AZs
Set group size:
Minimum: 2
Desired: 2
Maximum: 5
Health check type: ELB
Health check grace period: 60 seconds

Step 3 — Attach a Target Tracking Policy

◈ DIAGRAM
In the ASG → Automatic Scaling → Add Policy
Policy type: Target tracking
Metric: Average CPU utilization
Target value: 50%

This creates CloudWatch alarms automatically. No extra work needed.

Step 4 — Verify both instances are healthy

◈ DIAGRAM
EC2 → Auto Scaling Groups → devops-web-asg → Instance Management
Both instances should show: Lifecycle = InService, Health = Healthy

Step 5 — Test self-healing

◈ DIAGRAM
EC2 → Instances → select one of the ASG instances → Stop Instance
Wait 60 seconds then check:
EC2 → Auto Scaling Groups → devops-web-asg → Activity
You will see:
"Terminating EC2 instance: i-0abc123..."
"Launching a new EC2 instance..."
New instance registers and passes health check

Step 6 — Generate CPU load and watch scale-out

◈ DIAGRAM
SSH into one of the running instances:
ssh -i your-key.pem ec2-user@INSTANCE-IP
Run the stress tool for 5 minutes:
stress --cpu 4 --timeout 300
In the console, watch:
EC2 → Auto Scaling Groups → devops-web-asg → Monitoring → CPU Utilization
After 2-3 minutes CPU exceeds 50%. The Target Tracking alarm fires.
ASG launches more instances. Watch Activity History update in real time.
When stress ends, CPU drops. After the cooldown period ASG scales back in.

Step 7 — Cleanup

◈ DIAGRAM
EC2 → Auto Scaling Groups → devops-web-asg → Delete
Check "Force delete" to terminate all instances immediately
EC2 → Launch Templates → devops-web-lt → Delete

Common Mistakes to Avoid

Common Mistake

Setting minimum to 1. One instance means one Availability Zone. If that AZ has hardware trouble your app goes down until ASG replaces the instance — which takes minutes. Always set minimum to at least 2 across 2 AZs. High availability requires at least 2 copies.

Common Mistake

Using EC2 health check type instead of ELB. An instance can be running perfectly but your application crashing on every request. EC2 health check calls it healthy. ELB health check calls it unhealthy and ASG replaces it. Always use ELB health check for application workloads.

Common Mistake

Using a plain Amazon Linux AMI when you could pre-bake. Every 8-minute boot time is 8 minutes your new instances cannot serve traffic. Pre-bake your app. Boot in 90 seconds. Cooldown drops from 300 to 90 seconds. ASG reacts 3x faster to real spikes.

Security

Set Maximum to a number your budget can handle. During a traffic spike or a bug causing infinite scaling, the Maximum capacity is the only thing protecting your bill. Calculate your maximum based on cost, not just technical need.

Tip

For queue-based workers reading from SQS — scale on ApproximateNumberOfMessages not CPU. Set a target of messages-per-instance (say 1000 per worker). Queue builds up → instances scale out. Queue drains → instances scale in. Perfect matching between work and compute. CPU is irrelevant for a queue consumer.

Instance Refresh — Rolling Updates Without Downtime

When you update your Launch Template — new AMI, new instance type, new User Data — existing instances are still running the old version. Instance Refresh replaces them gradually without taking your application offline.

◈ DIAGRAM
You update Launch Template to v2 (new AMI with latest app)
↓
Trigger Instance Refresh with 90% minimum healthy threshold
↓
ASG replaces one instance at a time
Waits for new instance to pass health check before touching the next
↓
Eventually all instances running new Launch Template version
↓
Zero downtime. No manual work.
EC2 → Auto Scaling Groups → devops-web-asg → Instance Refresh → Start
Tip

Always trigger an Instance Refresh after updating a Launch Template. Without it, old instances keep running the old version indefinitely. Only new scale-out instances use the new template.

Resources

AWS Direct Connect vs Site-to-Site VPN Failover

AWS Direct Connect vs Site-to-Site VPN Failover

Direct Connect vs VPN isn't really either/or for production — it's a primary-plus-failover pattern. Here's how to design it, and when either/or is right.

5 min read•Aug 2026
Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot, at the crossover where Lambda stops being cheaper — 2026 pricing, invocation thresholds, and interruption math.

5 min read•Aug 2026
Secrets Manager vs Parameter Store vs Vault

Secrets Manager vs Parameter Store vs Vault

AWS Secrets Manager, Parameter Store, and HashiCorp Vault compared for 2026 - cost math, rotation, multi-cloud fit, and the Vault-to-OpenBao fork.

5 min read•Aug 2026
AWS VPC Security: Hardening Every Layer

AWS VPC Security: Hardening Every Layer

Most cloud security incidents start with a misconfigured VPC. Here's how to harden every layer — subnets, Security Groups, NACLs, and IAM — for production.

5 min read•Jul 2026
Event-Driven Architecture on AWS Explained

Event-Driven Architecture on AWS Explained

Event-driven architecture on AWS decouples services and absorbs traffic spikes using SQS, SNS, EventBridge, and Lambda — workflows that scale themselves.

5 min read•Jul 2026
S3 vs RDS vs DynamoDB: Choosing AWS Storage

S3 vs RDS vs DynamoDB: Choosing AWS Storage

Choosing S3, RDS, or DynamoDB wrong costs you in performance, cost, and scalability. Here is a practical decision guide based on your actual access patterns.

5 min read•Jul 2026
AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS bills surprise teams every month. Here are the 8 concrete actions that cut cloud spend by 40-60% without touching your application architecture.

5 min read•Jul 2026
EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2, Lambda, or Fargate — choosing the wrong AWS compute option costs you money and performance. Here is exactly when to use each one in production.

5 min read•Jul 2026

Explore More in AWS Compute and Auto Scaling

All 6 Topics

Frequently Asked Questions

Is Auto Scaling Groups - Dynamic Scaling for Production Workloads free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Auto Scaling Groups - Dynamic Scaling for Production Workloads topic cover?

Configure Auto Scaling Groups with Launch Templates, scaling policies, and CloudWatch alarms to automatically adjust EC2 fleet size based on real demand.