Auto Scaling Group
An Auto Scaling Group (ASG) manages a fleet of EC2 instances automatically — launching new instances when load increases, terminating them when load drops, and replacing unhealthy ones without human involvement. It operates between a defined minimum and maximum instance count, adjusting toward a desired count driven by scaling policies such as target tracking or scheduled scaling.
Zerodha scales its API fleet on a schedule tied to market hours — desired capacity rises automatically at 8:30 AM before market open and scales back down after 4 PM close, using a Scheduled scaling policy layered on top of Target Tracking for intraday spikes.
How ASG and the Load Balancer Cooperate
When an instance fails its load balancer health check, the LB stops routing to it, the ASG terminates it, and launches a fresh replacement from the Launch Template — typically a 3-5 minute self-healing loop with zero human action.
Cooldown Matters
After scaling, the ASG waits a cooldown period before scaling again to avoid over-provisioning while new instances are still booting. A pre-baked custom AMI that boots in 90 seconds lets you set a much shorter cooldown, so the ASG reacts faster to real spikes.
RememberThe ASG itself is free — you only pay for the EC2 instances it creates, so there's no cost reason to skip it even for workloads you don't expect to scale.
Frequently Asked Questions
What's the difference between an ASG's minimum, maximum, and desired capacity?
Minimum and maximum set the hard bounds the ASG will never go below or above, regardless of scaling signals. Desired capacity is the actual target the ASG is currently trying to maintain, adjusted dynamically by scaling policies (like target tracking on CPU utilization) within those bounds. If an instance fails a health check, the ASG replaces it to keep the running count at desired capacity — that's the self-healing behavior separate from scaling up or down.
What's a common mistake when configuring an ASG?
Setting scaling policies based only on CPU utilization when the actual bottleneck is something else — memory, connection count, or request latency — so the ASG scales confidently in the wrong direction while the real problem persists. Also common: forgetting that abrupt scale-in during a deploy can terminate instances mid-request if connection draining isn't configured, causing dropped requests during otherwise routine scaling events.