Skip to main content

AWS CloudWatch Monitoring and Observability

Learn AWS CloudWatch monitoring with metrics, alarms, logs, CloudTrail, Config and EventBridge, and set up alerts that do not cause alert fatigue.

~3 hours
11 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Observability Is Several Separate Problems

The 2 AM incident that needed four different tools It is 2 AM.

Collecting Metrics and Custom Metrics

How metrics, namespaces, and dimensions fit together A metric is a number tracked over time, such as CPU utilization or queue depth.

Building Dashboards and Alarms That Do Not Cause Alert Fatigue

Dashboards that answer the first question fast A dashboard is one screen showing what matters for one system, so the on-call engineer does not hunt...

Working with CloudWatch Logs, Logs Insights, and Metric Filters

Log groups, log streams, and retention CloudWatch Logs stores logs from your applications and services.

Collecting Memory and Disk Metrics with the CloudWatch Agent

Why EC2 cannot report memory by itself AWS sees the hypervisor, which means CPU and network.

Tracing Requests with X-Ray and Application Signals

Why metrics and logs cannot show where the time went Metrics tell you that error rate went up. Logs tell you which request threw which exception.

Skills You'll Master

CLOUDWATCHCLOUDTRAILAWS-CONFIGEVENTBRIDGEMONITORING

Curriculum Index11 topics

1

Understanding Why Observability Is Several Separate Problems

The 2 AM incident that needed four different tools It is 2 AM.

2

Collecting Metrics and Custom Metrics

How metrics, namespaces, and dimensions fit together A metric is a number tracked over time, such as CPU utilization or...

3

Building Dashboards and Alarms That Do Not Cause Alert Fatigue

Dashboards that answer the first question fast A dashboard is one screen showing what matters for one system, so the...

4

Working with CloudWatch Logs, Logs Insights, and Metric Filters

Log groups, log streams, and retention CloudWatch Logs stores logs from your applications and services.

5

Collecting Memory and Disk Metrics with the CloudWatch Agent

Why EC2 cannot report memory by itself AWS sees the hypervisor, which means CPU and network.

6

Tracing Requests with X-Ray and Application Signals

Why metrics and logs cannot show where the time went Metrics tell you that error rate went up.

7

Auditing API Calls with CloudTrail

What CloudTrail records CloudTrail records API calls made in your account: console clicks, CLI commands, SDK calls, and...

8

Tracking Configuration Drift with AWS Config

What Config answers that CloudTrail and CloudWatch cannot AWS Config records how resources are configured over time and...

9

Reacting in Real Time with EventBridge

How events flow from AWS actions to automation CloudWatch says something is unhealthy. CloudTrail says who acted.

10

Hands-on Lab: Building an Observability and Governance Pipeline

Before you start: cost and prerequisites This lab builds all four tools around one small instance.

11

Quick Reference and Common Mistakes

Quick reference Common mistakes Looking for memory under AWS/EC2 is the most common monitoring mistake in the first...

Career Impact

Roles that use the skills in this module.

  • Site Reliability Engineer

  • DevOps Engineer

  • Platform Engineer

  • Cloud Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

CloudWatch tells you how your systems are behaving right now through metrics, logs, and alarms. CloudTrail records who called which AWS API and when. You need CloudWatch to notice a problem and CloudTrail to find out who or what caused it.

AWS can only see the hypervisor side of an instance, such as CPU and network. Memory lives inside your operating system, so you must install the CloudWatch agent to publish it as a custom metric under the CWAgent namespace.

No. Config is a detective control: it records changes and flags non-compliant resources after the change has happened. To block an action you need IAM policies or Service Control Policies.

Event history in the console keeps 90 days of management events at no charge. For anything longer, or for data events, you create a trail that delivers logs to S3 or use CloudTrail Lake.