Skip to main content

VM Scale Set

An Azure Virtual Machine Scale Set (VMSS) manages a group of identical, load-balanced VMs that scale automatically based on demand or a defined schedule, and automatically replaces unhealthy instances. It's the Azure equivalent of an AWS Auto Scaling Group — a single model defines the VM image, size, and configuration applied uniformly across every instance in the set.

VM Scale Set

A Scale Set manages a fleet of identical VMs as a single resource, handling creation, load balancing, and autoscaling based on rules you define.

Why It Matters in Production

Swiggy runs its order-matching service on a VM Scale Set that autoscales from 10 to 80 instances during the dinner-hour rush, then scales back down overnight to control cost.

az vmss create --resource-group swiggy-prod-rg --name order-matching-vmss
--image Ubuntu2204 --instance-count 10 --vm-sku Standard_D4s_v5
--load-balancer swiggy-lb

Tip

Use a custom VM image with your app pre-baked in the Scale Set to cut instance boot-to-ready time during scale-out events.

Frequently Asked Questions

How does a VM Scale Set decide when to add or remove instances?

VMSS uses autoscale rules tied to Azure Monitor metrics — CPU percentage, memory, queue length, or custom metrics — with defined thresholds, cooldown periods, and min/max instance counts. You can also scale on a fixed schedule (e.g. add capacity at 9am weekdays). Unlike manually managed VMs, every instance in the set is created from one shared model definition, so scaling out clones an identical, already-configured VM rather than requiring per-instance setup.

What's a common mistake when configuring VMSS autoscale rules?

Setting scale-out and scale-in thresholds too close together (e.g. scale out above 70% CPU, scale in below 65%) causes flapping — the set repeatedly adds and removes instances as load hovers near the boundary, generating churn and unnecessary billing. Leave a meaningful gap between thresholds and set a sensible cooldown period, and remember that scale-in can terminate instances mid-request unless you configure instance protection or graceful termination for in-flight connections.