Overview and What You Will Learn
In this lab, you will create a Notification Channel, then create an alert policy referencing it, and confirm the full path from a metric crossing a threshold to an actual email arriving - the same connected chain every cloud platform's monitoring system needs to genuinely protect a production workload.
Why This Matters in Production
A team configures an alert policy watching disk space, but never attaches a Notification Channel to it, assuming the alert "exists" and therefore someone will be told. The alert fires correctly and logs its own history in the Monitoring console the entire time - but nobody is ever actually notified, and the disk fills up completely before anyone discovers the problem, days after the alert had already been silently firing.
Core Principles
An alert policy has two genuinely separate parts, and both must be configured correctly for the alert to actually protect anything.
+------------------------------------------+| Metric crosses the configured threshold |+------------------------------------------+ | v+------------------------------------------+| Alert Policy fires || (this always happens if the condition is || met, and is always logged) |+------------------------------------------+ | v+------------------------------------------+| Notification Channel || (email, SMS, PagerDuty, webhook) || Without one attached, nobody is told |+------------------------------------------+Detailed Step-by-Step Practical Lab
- Create a project and a test VM to monitor:
gcloud projects create gcp-monitoring-lab-2026 --name="Monitoring Lab"gcloud config set project gcp-monitoring-lab-2026gcloud services enable compute.googleapis.com monitoring.googleapis.com gcloud compute instances create vm-monitor-test \ --zone=asia-south1-a \ --machine-type=e2-small \ --image-family=debian-12 \ --image-project=debian-cloud- Create a Notification Channel for email:
gcloud alpha monitoring channels create \ --display-name="Ops Team Email" \ --type=email \ --channel-labels=email_address=ops-team@example.com- Retrieve the channel's ID for use in the alert policy:
gcloud alpha monitoring channels list \ --filter="displayName='Ops Team Email'" \ --format="value(name)"- Create an alert policy watching CPU utilization, explicitly referencing the Notification Channel:
gcloud alpha monitoring policies create \ --notification-channels=CHANNEL_ID_FROM_STEP_3 \ --display-name="High CPU Alert" \ --condition-display-name="CPU above 85%" \ --condition-filter='resource.type="gce_instance" AND metric.type="compute.googleapis.com/instance/cpu/utilization"' \ --condition-threshold-value=0.85 \ --condition-threshold-duration=300s- Confirm the policy was created with the channel correctly attached:
gcloud alpha monitoring policies list \ --format="value(displayName,notificationChannels)"- Deliberately create a test alert with no channel attached, to observe the exact silent-failure mode this lab is warning about:
gcloud alpha monitoring policies create \ --display-name="Silent Test Alert (No Channel)" \ --condition-display-name="CPU above 85%" \ --condition-filter='resource.type="gce_instance" AND metric.type="compute.googleapis.com/instance/cpu/utilization"' \ --condition-threshold-value=0.85 \ --condition-threshold-duration=300sNoteThis second policy will still evaluate and log its own firing history in the Monitoring console exactly like the first one - the only difference is that nobody will actually receive a notification when it fires, since no channel was attached. This is precisely the failure mode described at the start of this lab.
- Clean up:
gcloud alpha monitoring policies list --format="value(name)" | \ while read policy; do gcloud alpha monitoring policies delete "$policy" --quiet; donegcloud compute instances delete vm-monitor-test --zone=asia-south1-a --quietgcloud projects delete gcp-monitoring-lab-2026 --quietProduction Best Practices & Common Pitfalls
Common MistakeCreating an alert policy and assuming it's complete without verifying a Notification Channel is actually attached to it. The policy will fire and log correctly either way - the only symptom of a missing channel is silence, which is easy to miss until an actual incident reveals nobody was ever told.
TipSet the condition's evaluation window (
--condition-threshold-duration) to a realistic value for the metric's normal behavior - 300 seconds (5 minutes) or more for something naturally bursty like CPU, to avoid brief, harmless spikes triggering repeated false alarms.
- Route different severities to different Notification Channels. A critical production outage alert going to an on-call phone number via SMS or PagerDuty is a different urgency than a minor cost-threshold warning going to a shared team email - treating every alert identically dilutes the ones that actually need immediate attention.
- A Dashboard is a separate concept from an alert policy - dashboards visualize metrics for humans actively looking at them, while alert policies proactively notify without anyone needing to be watching in the first place. Both matter, but they solve different problems.
Quick Reference & Troubleshooting Commands
| Command | Description |
|---|---|
gcloud alpha monitoring channels create |
Create a new Notification Channel |
gcloud alpha monitoring policies create --notification-channels= |
Create an alert policy with a channel attached |
gcloud alpha monitoring policies list |
List existing alert policies |
gcloud alpha monitoring channels list |
List existing Notification Channels |