Skip to main content

CloudWatch Alarm

A CloudWatch Alarm watches a single metric and automatically triggers an action — stopping or recovering an EC2 instance, scaling an Auto Scaling Group, or notifying via SNS — when the metric crosses a defined threshold. Alarms exist in one of three states: OK, ALARM, or INSUFFICIENT_DATA, and enable automated operational responses without requiring a human to be watching a dashboard.

An SRE team wires a CloudWatch Alarm on StatusCheckFailed_System to automatically trigger EC2 Instance Recovery — when underlying hardware fails, AWS migrates the instance to healthy hardware while preserving its private IP, Elastic IP, and instance ID, with zero manual intervention.

Composite Alarms Reduce Alert Fatigue

A single CPU-threshold alarm firing on every brief spike trains engineers to ignore pages. A Composite Alarm combining "CPU > 80% AND Memory > 90%, both for 5 minutes" only fires when both conditions are genuinely true together, cutting false positives sharply.

Tip

Always test a new alarm by forcing it into ALARM state manually (aws cloudwatch set-alarm-state) to confirm the downstream SNS notification, scaling action, or Lambda actually fires — an untested alarm often silently fails to work when it's needed most.

Frequently Asked Questions

Why does a CloudWatch Alarm have an INSUFFICIENT_DATA state instead of just OK or ALARM?

INSUFFICIENT_DATA means the alarm hasn't yet received enough data points to evaluate the threshold — common right after an alarm is created, or when a metric source (like a stopped EC2 instance) stops publishing data entirely. Treating this state as equivalent to OK is a mistake: it typically means monitoring itself has a gap, which is arguably worse than a known ALARM state because you have no visibility into whether the underlying condition is fine or broken.

What's a common mistake when configuring CloudWatch Alarms?

Setting the evaluation period and datapoints-to-alarm too aggressively (e.g., 1 datapoint over 1 minute), which triggers on brief, self-resolving spikes and trains the team to ignore alerts. A steadier pattern — like 3 out of 5 one-minute periods breaching the threshold — filters noise while still catching real problems quickly. Also easy to forget: an alarm with no configured action (SNS topic, Auto Scaling policy) just sits in the console changing color with nobody watching it.