Overview and What You Will Learn
In this lab, you will create an Action Group defining who gets notified, then build an alert rule that watches a VM's CPU metric and triggers that Action Group once a threshold is crossed - the complete path from a raw metric to an actual notification landing in someone's inbox.
Why This Matters in Production
A support engineer at CRED gets paged at 2 AM because a production database ran out of disk space - an event a properly configured Azure Monitor alert would have flagged hours earlier, giving the team time to react before customers noticed anything. The gap between "the metric existed" and "someone was actually notified in time" is exactly what Action Groups and alert rules close.
Core Principles
Azure Monitor alerts have two separate parts that must both be configured correctly: the condition that triggers the alert, and the Action Group defining what happens once it does.
+------------------------------------------+| Metric collected continuously || (e.g. Percentage CPU on a VM) |+------------------------------------------+ | v+------------------------------------------+| Alert Rule evaluates the condition || (e.g. average > 85% for 5 minutes) |+------------------------------------------+ | v+------------------------------------------+| Action Group fires || (email, SMS, webhook, or trigger automation)|+------------------------------------------+An alert rule with no Action Group attached will still show up in the Azure Portal's alert history, but nobody will actually be notified - the Action Group is what makes an alert actionable rather than just a passive log entry.
Detailed Step-by-Step Practical Lab
- Create a Resource Group and a test VM to monitor:
az group create --name rg-monitor-lab-mumbai --location centralindia az vm create \ --resource-group rg-monitor-lab-mumbai \ --name vm-monitor-lab \ --image Ubuntu2204 \ --size Standard_B1s \ --admin-username azureadmin \ --generate-ssh-keys- Create an Action Group defining an email recipient to notify:
az monitor action-group create \ --resource-group rg-monitor-lab-mumbai \ --name ag-oncall-team \ --short-name oncall \ --action email oncall-engineer rahul@example.com- Create an alert rule watching CPU usage, attached to that Action Group:
VM_ID=$(az vm show \ --resource-group rg-monitor-lab-mumbai \ --name vm-monitor-lab \ --query "id" --output tsv) az monitor metrics alert create \ --name "high-cpu-alert-lab" \ --resource-group rg-monitor-lab-mumbai \ --scopes "$VM_ID" \ --condition "avg Percentage CPU > 85" \ --window-size 5m \ --evaluation-frequency 1m \ --action ag-oncall-team- Confirm the alert rule was created and check its current status:
az monitor metrics alert show \ --resource-group rg-monitor-lab-mumbai \ --name "high-cpu-alert-lab" \ --query "enabled" --output tsvSimulate load on the VM (using a tool like
stressinstalled via SSH) to trigger the alert, then check the Action Group's configured email address for the notification.Review the alert's history in the Portal or via CLI to confirm it fired as expected:
az monitor activity-log alert list \ --resource-group rg-monitor-lab-mumbai \ --output table- Clean up:
az group delete --name rg-monitor-lab-mumbai --yes --no-waitProduction Best Practices & Common Pitfalls
Common MistakeCreating an alert rule but forgetting to attach an Action Group, or attaching one with an outdated email address nobody actually monitors. The alert will still technically "fire" and appear in Azure's own alert history, but if nobody is actually notified in a way they'll see promptly, the alert provides no real protection.
TipAlert on the metric that predicts a problem, not only the metric that confirms one already happened. A disk-space-approaching-full alert gives you time to react before an outage; an "application crashed" alert means the outage has already started.
- Set the evaluation window and frequency thoughtfully. A window too short can trigger false alarms from brief, harmless spikes; a window too long delays a genuine alert past the point where it's still useful.
- Route different severities to different Action Groups. A critical production outage alert going to an on-call phone number is a different urgency than a minor cost-threshold warning going to a shared team email - treating every alert identically dilutes the ones that actually need immediate attention.
Quick Reference & Troubleshooting Commands
| Command | Description |
|---|---|
az monitor action-group create |
Create a new Action Group |
az monitor metrics alert create |
Create a metric-based alert rule |
az monitor metrics alert show |
Check an alert rule's configuration and status |
az monitor activity-log alert list |
Review alert firing history |