Overview and What You Will Learn
In this lab, you will create a Virtual Machine Scale Set, configure an autoscale rule that adds instances when CPU load rises, and simulate load to watch the scale set actually add capacity automatically, then remove it again once load drops.
Why This Matters in Production
During a flash sale, a CRED-style fintech app needs to handle a sudden ten-fold spike in traffic without the engineering team manually launching new VMs at 2 AM, and without paying for that same peak capacity to sit idle for the other 364 days of the year. A properly configured scale set handles both sides of this problem automatically.
Core Principles
A Virtual Machine Scale Set manages a group of identical, load-balanced VMs as a single unit. Instead of manually deciding when to add or remove capacity, an autoscale rule watches a metric and adjusts instance count on its own.
+------------------------------------------+| CPU metric monitored continuously |+------------------------------------------+ | v+------------------------------------------+| Average CPU > threshold for X minutes? |+------------------------------------------+ | | yes no | | v v+------------------+ +------------------+| Scale OUT | | Scale IN || (add instances) | | (remove instances) |+------------------+ +------------------+The scale set's min, max, and current instance count define the boundaries autoscaling operates within - it will never scale below the minimum or above the maximum, regardless of how extreme the metric readings become.
Detailed Step-by-Step Practical Lab
- Create a Resource Group:
az group create --name rg-vmss-lab-mumbai --location centralindia- Create a Virtual Machine Scale Set with an initial instance count of 2:
az vmss create \ --resource-group rg-vmss-lab-mumbai \ --name vmss-web-lab \ --image Ubuntu2204 \ --instance-count 2 \ --admin-username azureadmin \ --generate-ssh-keys \ --vm-sku Standard_B1s- Confirm both instances are running:
az vmss list-instances \ --resource-group rg-vmss-lab-mumbai \ --name vmss-web-lab \ --output table- Create an autoscale profile with a minimum of 2 and a maximum of 6 instances:
az monitor autoscale create \ --resource-group rg-vmss-lab-mumbai \ --resource vmss-web-lab \ --resource-type Microsoft.Compute/virtualMachineScaleSets \ --name autoscale-web-lab \ --min-count 2 \ --max-count 6 \ --count 2- Add a scale-out rule that adds one instance when average CPU exceeds 70% for 5 minutes:
az monitor autoscale rule create \ --resource-group rg-vmss-lab-mumbai \ --autoscale-name autoscale-web-lab \ --condition "Percentage CPU > 70 avg 5m" \ --scale out 1- Add a scale-in rule that removes one instance when average CPU drops below 30% for 5 minutes:
az monitor autoscale rule create \ --resource-group rg-vmss-lab-mumbai \ --autoscale-name autoscale-web-lab \ --condition "Percentage CPU < 30 avg 5m" \ --scale in 1NoteScale-out and scale-in rules are configured independently and often use different thresholds on purpose. A wider gap between the two (like 70% out, 30% in) prevents the scale set from rapidly oscillating instances up and down when load hovers near a single shared threshold.
- Simulate CPU load on the instances (using a tool like
stressinstalled on the VM) and watch the instance count grow:
az vmss list-instances \ --resource-group rg-vmss-lab-mumbai \ --name vmss-web-lab \ --output table## Re-run this periodically during the load test to see new instances appear- Clean up:
az group delete --name rg-vmss-lab-mumbai --yes --no-waitProduction Best Practices & Common Pitfalls
Common MistakeSetting the scale-out and scale-in thresholds too close together, like 60% out and 55% in. When load hovers naturally around that range, the scale set can rapidly add and remove instances back and forth, wasting time and money on constant scaling activity instead of settling into a stable state.
TipSet a Cooldown period on autoscale rules so the scale set waits a reasonable amount of time after a scaling action before evaluating whether to scale again. This prevents overly aggressive back-to-back scaling decisions based on brief, temporary metric spikes.
- Scale sets update by rolling out a new VM image, not by patching running instances individually. Build your application updates into a new image version and let the scale set roll it out, rather than trying to manually update every running instance.
- Instance count boundaries (min/max) should reflect real cost and capacity planning, not arbitrary round numbers. A max count set too low can silently cap your application's ability to handle a genuine traffic spike.
Quick Reference & Troubleshooting Commands
| Command | Description |
|---|---|
az vmss create |
Create a new Virtual Machine Scale Set |
az vmss list-instances |
List current instances in a scale set |
az monitor autoscale create |
Create an autoscale profile |
az monitor autoscale rule create |
Add a scale-out or scale-in rule |