AWS CloudWatch Monitoring and Observability
Learn AWS CloudWatch monitoring with metrics, alarms, logs, CloudTrail, Config and EventBridge, and set up alerts that do not cause alert fatigue.
What You'll Learn
Understanding Why Observability Is Several Separate Problems
The 2 AM incident that needed four different tools It is 2 AM.
Collecting Metrics and Custom Metrics
How metrics, namespaces, and dimensions fit together A metric is a number tracked over time, such as CPU utilization or queue depth.
Building Dashboards and Alarms That Do Not Cause Alert Fatigue
Dashboards that answer the first question fast A dashboard is one screen showing what matters for one system, so the on-call engineer does not hunt...
Working with CloudWatch Logs, Logs Insights, and Metric Filters
Log groups, log streams, and retention CloudWatch Logs stores logs from your applications and services.
Collecting Memory and Disk Metrics with the CloudWatch Agent
Why EC2 cannot report memory by itself AWS sees the hypervisor, which means CPU and network.
Tracing Requests with X-Ray and Application Signals
Why metrics and logs cannot show where the time went Metrics tell you that error rate went up. Logs tell you which request threw which exception.
Skills You'll Master
Curriculum Index11 topics
Understanding Why Observability Is Several Separate Problems
The 2 AM incident that needed four different tools It is 2 AM.
Collecting Metrics and Custom Metrics
How metrics, namespaces, and dimensions fit together A metric is a number tracked over time, such as CPU utilization or...
Building Dashboards and Alarms That Do Not Cause Alert Fatigue
Dashboards that answer the first question fast A dashboard is one screen showing what matters for one system, so the...
Working with CloudWatch Logs, Logs Insights, and Metric Filters
Log groups, log streams, and retention CloudWatch Logs stores logs from your applications and services.
Collecting Memory and Disk Metrics with the CloudWatch Agent
Why EC2 cannot report memory by itself AWS sees the hypervisor, which means CPU and network.
Tracing Requests with X-Ray and Application Signals
Why metrics and logs cannot show where the time went Metrics tell you that error rate went up.
Auditing API Calls with CloudTrail
What CloudTrail records CloudTrail records API calls made in your account: console clicks, CLI commands, SDK calls, and...
Tracking Configuration Drift with AWS Config
What Config answers that CloudTrail and CloudWatch cannot AWS Config records how resources are configured over time and...
Reacting in Real Time with EventBridge
How events flow from AWS actions to automation CloudWatch says something is unhealthy. CloudTrail says who acted.
Hands-on Lab: Building an Observability and Governance Pipeline
Before you start: cost and prerequisites This lab builds all four tools around one small instance.
Quick Reference and Common Mistakes
Quick reference Common mistakes Looking for memory under AWS/EC2 is the most common monitoring mistake in the first...
Career Impact
Roles that use the skills in this module.
Site Reliability Engineer
DevOps Engineer
Platform Engineer
Cloud Engineer
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
CloudWatch tells you how your systems are behaving right now through metrics, logs, and alarms. CloudTrail records who called which AWS API and when. You need CloudWatch to notice a problem and CloudTrail to find out who or what caused it.
AWS can only see the hypervisor side of an instance, such as CPU and network. Memory lives inside your operating system, so you must install the CloudWatch agent to publish it as a custom metric under the CWAgent namespace.
No. Config is a detective control: it records changes and flags non-compliant resources after the change has happened. To block an action you need IAM policies or Service Control Policies.
Event history in the console keeps 90 days of management events at no charge. For anything longer, or for data events, you create a trail that delivers logs to S3 or use CloudTrail Lake.