Skip to main content

CloudWatch, CloudTrail, and Config - Monitoring, Auditing, and Compliance

Monitor AWS infrastructure with CloudWatch metrics and alarms, audit every API call with CloudTrail, and enforce compliance rules with AWS Config.

What you will learn

  • CloudWatch Metrics — namespaces, dimensions, and why RAM is not collected by default
  • Custom Metrics — pushing your own business data into CloudWatch
  • CloudWatch Logs — log groups, streams, and the Unified Agent
  • Subscription Filters — shipping logs in real time to Lambda, Kinesis, or Firehose
  • CloudWatch Logs Insights — querying your logs like a database
  • CloudWatch Alarms — states, composite alarms, and EC2 auto-recovery
  • The four CloudWatch Insights features — Container, Lambda, Contributor, Application
  • Amazon EventBridge — scheduling tasks and reacting to AWS events automatically
  • AWS CloudTrail — who did what, when, and from where
  • AWS Config — are my resources configured correctly right now
  • CloudWatch vs CloudTrail vs Config — which service answers which question

Why this matters

A Zerodha trading platform going down at market open. A Razorpay payment API throwing 500 errors. An EC2 instance being terminated by an unknown IAM user at 2 AM. None of these can be investigated without the right visibility tools in place.

CloudWatch tells you something is wrong right now. CloudTrail tells you who did what and when. Config tells you whether your infrastructure drifted from its intended state. Together they form the observability and compliance backbone of every serious AWS deployment. Without them you are flying blind — debugging by guesswork instead of data.

CloudWatch Metrics — What AWS Measures Automatically

CloudWatch collects metrics from every AWS service automatically. A metric is a numeric value tracked over time — CPU utilisation, network bytes, request count, disk reads.

Key concepts:

◈ DIAGRAM
Namespace → the folder that groups metrics per service
EC2 metrics live in: AWS/EC2
Lambda metrics live in: AWS/Lambda
RDS metrics live in: AWS/RDS
Dimension → extra context that identifies which specific resource
InstanceId = i-0abc123 (which EC2)
FunctionName = devops-orders (which Lambda)
DBInstanceIdentifier = prod-db (which RDS)
Period → how often the metric is evaluated
Standard: every 60 seconds
High resolution: every 1, 5, 10, or 30 seconds
Remember

RAM (memory usage) is NOT collected by EC2 by default. CloudWatch gets CPU, network, disk I/O, and status checks from EC2 automatically — but memory requires the CloudWatch Unified Agent installed on the instance. This is one of the most commonly tested CloudWatch facts.

Custom Metrics — push your own data:

Any number you care about can become a CloudWatch metric. Active users, queue depth, business transaction count, cache hit rate — anything your application can measure.

Bash
## Push a custom metric — active sessions count
aws cloudwatch put-metric-data \
--namespace "DevOpsNetwork/App" \
--metric-name "ActiveSessions" \
--value 1247 \
--unit Count \
--dimensions Environment=prod,Service=orders \
--region ap-south-1

Custom metrics appear in CloudWatch alongside AWS metrics and can trigger alarms just like any built-in metric.

CloudWatch Logs — Centralised Log Storage

CloudWatch Logs stores application logs, infrastructure logs, and audit logs from across your entire AWS environment.

Core structure:

Bash
Log Group → one per application or service
/aws/lambda/devops-orders
/aws/rds/instance/prod-db/error
/vpc/prod-vpc-flow-logs
Log Stream → logs from one specific source inside the group
A new stream per Lambda execution environment
One stream per EC2 instance
Retention → how long logs are kept
Default: never expire (costs money over time)
Set explicitly: 1 day to 10 years, or never expire
Best practice: set retention on every log group

What sends logs to CloudWatch automatically:

◈ DIAGRAM
Lambda → all print() and console.log() output
API Gateway → access logs (when enabled)
VPC Flow Logs → network traffic records
CloudTrail → API call audit events
RDS → error logs, slow query logs (when enabled)

EC2 sends NO logs by default. You must install the CloudWatch Agent.

Logs Agent vs Unified Agent:

Logs Agent Unified Agent
Version Old New — always use this
Sends logs Yes Yes
Sends system metrics No Yes — RAM, CPU detail, disk, netstat
Config management Manual config file Via SSM Parameter Store (centralised)

Always use the Unified Agent. It does everything the old agent does plus gives you RAM usage, process counts, disk usage, TCP connections — metrics EC2 does not expose natively.

Where logs can go from CloudWatch:

◈ DIAGRAM
CloudWatch Logs
├── S3 (archiving — up to 12 hour delay, use for long-term storage)
├── Kinesis Data Streams (real-time to your own consumers)
├── Kinesis Data Firehose (near real-time to S3, Redshift, OpenSearch)
└── Lambda (trigger a function when a log pattern matches)
Remember

Exporting to S3 (CreateExportTask) takes up to 12 hours. It is for archiving, not real-time alerting. For real-time log processing use Subscription Filters — they deliver instantly.

Subscription Filters — Real-Time Log Routing

A Subscription Filter watches a log group and forwards matching log events in real time to a destination.

Bash
/aws/lambda/devops-orders → filter pattern: "ERROR"
↓
Only ERROR log lines forwarded → Lambda function
↓
Lambda sends Slack alert immediately

Multi-account log aggregation:

Production, staging, and dev accounts each forward logs via Subscription Filters to a central Kinesis Data Streams in a security account → Kinesis Firehose → centralised S3 bucket. One bucket, all logs, searchable with Athena.

CloudWatch Logs Insights — Query Your Logs

Logs Insights lets you run SQL-like queries against your stored log data. Filter, aggregate, sort, count — across one or multiple log groups.

TEXT
## Find all errors in the last hour
fields @timestamp, @message
| filter @message like /ERROR/
| sort @timestamp desc
| limit 50
## Count errors per 5-minute window
fields @timestamp, @message
| filter @message like /Exception/
| stats count(*) as errors by bin(5m)
## Find slowest Lambda invocations
filter @type = "REPORT"
| stats avg(@duration), max(@duration) by bin(5m)
| sort max(@duration) desc

Results can be saved as named queries and pinned directly to a CloudWatch Dashboard.

Remember

Logs Insights queries historical data already in CloudWatch. It is not real-time streaming. For real-time log processing use Subscription Filters + Lambda.

CloudWatch Alarms — Act When Something Goes Wrong

An alarm watches a single metric and fires when it crosses a threshold you define.

Three states:

◈ DIAGRAM
OK → metric is within acceptable range
ALARM → threshold breached
INSUFFICIENT_DATA → not enough data points yet to evaluate

What an alarm can trigger:

◈ DIAGRAM
EC2 Action → stop, terminate, reboot, or recover the instance
Auto Scaling → scale out or scale in
SNS → send notification (email, Slack via Lambda, PagerDuty)
Bash
## Create alarm: fire when average CPU > 80% for 5 minutes
aws cloudwatch put-metric-alarm \
--alarm-name "high-cpu-prod" \
--metric-name CPUUtilization \
--namespace AWS/EC2 \
--dimensions Name=InstanceId,Value=i-0abc123 \
--statistic Average \
--period 300 \
--threshold 80 \
--comparison-operator GreaterThanThreshold \
--evaluation-periods 1 \
--alarm-actions arn:aws:sns:ap-south-1:123456789012:devops-alerts \
--region ap-south-1
## Test immediately without waiting for real CPU spike
aws cloudwatch set-alarm-state \
--alarm-name "high-cpu-prod" \
--state-value ALARM \
--state-reason "testing" \
--region ap-south-1

Composite Alarms — reduce alert fatigue:

A single alarm firing on every small CPU spike creates noise. Engineers start ignoring pages. A Composite Alarm combines multiple alarms with AND/OR logic.

◈ DIAGRAM
Alarm A (CPU > 80%) ─┐
├──→ Composite Alarm (CPU AND Memory both high) ──→ SNS page
Alarm B (Memory > 85%) ─┘
CPU spikes briefly alone → Composite stays OK → no page
CPU AND Memory both high → Composite fires → page the on-call engineer

Composite Alarms dramatically reduce alert noise. You only get paged when a real problem is confirmed.

EC2 Instance Recovery:

CloudWatch can automatically recover a broken EC2 instance when the underlying hardware fails.

◈ DIAGRAM
CloudWatch monitors EC2 status checks
↓
StatusCheckFailed_System alarm fires (hardware issue, not app issue)
↓
EC2 Instance Recovery triggered automatically
↓
Instance migrated to healthy hardware
Keeps: same private IP, same public IP, same Elastic IP, same metadata

No data loss. No human action. The instance recovers itself.

CloudWatch Insights — Four Specialised Features

Container Insights:

Collects metrics and logs from ECS, EKS, and Kubernetes. Shows cluster, node, pod, and task-level visibility in pre-built dashboards. For EKS, CloudWatch deploys a containerised agent inside the cluster automatically.

Lambda Insights:

Monitoring specifically for Lambda. Collects CPU time, memory, disk, network, cold start count, and worker shutdowns. Deployed as a Lambda Layer — attach to your function, no code changes needed.

Contributor Insights:

Reads log data and finds the top contributors to a metric. Who or what is causing the most traffic, errors, or load?

TEXT
Example: find the top 10 IPs generating the most 404 errors in ALB access logs
Example: find the top 5 DynamoDB partition keys being accessed most frequently

Application Insights:

Automated monitoring for your entire application stack. Point it at your EC2 application resources, it discovers related components (RDS, ELB, Auto Scaling), creates dashboards automatically, and sends findings when problems occur. Uses machine learning for anomaly detection.

Amazon EventBridge — React to Everything

EventBridge is an event router. Something happens anywhere in AWS — EC2 state changes, S3 uploads, CodeBuild failures, IAM actions — EventBridge sees it and routes it to a target.

Two ways to use it:

◈ DIAGRAM
Schedule → run something on a cron or rate-based timer
Every day at 2 AM → EventBridge → Lambda → delete old logs
Every 5 minutes → EventBridge → Lambda → check health of external API
Event Pattern → react when something specific happens
Root user signs in → EventBridge → SNS → email alert immediately
EC2 instance stopped → EventBridge → Lambda → send Slack notification
S3 object deleted → EventBridge → Lambda → log the deletion for audit

Event buses:

◈ DIAGRAM
Default Event Bus → all AWS service events (EC2, S3, Lambda, etc)
Partner Event Bus → third-party SaaS events (Datadog, Zendesk, Salesforce)
Custom Event Bus → your own application events

Schema Registry:

EventBridge automatically detects the structure of events flowing through it. Your application can download a schema and generate typed code that already knows the field names and types — no manual inspection needed.

Real production example:

Bash
IAM root user logs in at 2 AM
↓
CloudTrail records the event
↓
EventBridge rule matches: source=aws.signin, userIdentity.type=Root
↓
SNS sends email to security team immediately
↓
Lambda automatically disables the root access key (if one exists)

All of this happens automatically. Nobody manually watched CloudTrail.

AWS CloudTrail — The Audit Log

CloudTrail records every API call made in your AWS account — console clicks, CLI commands, SDK calls, and service-to-service calls. It answers: who did what, when, and from where.

TEXT
CloudTrail is enabled by default
Events stored free for 90 days in CloudTrail Event History
For longer retention: create a Trail to send events to S3

Three event types:

◈ DIAGRAM
Management Events (logged by default):
Operations performed on AWS resources
Examples: creating an EC2, modifying a Security Group, adding an IAM policy
Read and Write events can be separated to reduce noise
Data Events (NOT logged by default — high volume):
Object-level operations inside resources
Examples: S3 GetObject, PutObject, DeleteObject; Lambda Invoke
Must be explicitly enabled — extra cost
Critical for: compliance on sensitive buckets, auditing Lambda invocations
CloudTrail Insights Events:
Unusual activity detection using baseline analysis
Detects: unexpected API call spikes, unusual IAM action bursts
When anomaly found → stored in S3, visible in console, EventBridge event fired
Remember

S3 object-level actions (GetObject, PutObject, DeleteObject) are Data Events and are NOT logged by default. Teams regularly miss this when setting up compliance logging. Always explicitly enable Data Events for any S3 bucket containing sensitive data.

Storing trails long-term:

◈ DIAGRAM
CloudTrail → Trails → Create Trail
Trail name: prod-audit-trail
S3 bucket: create new → devops-cloudtrail-logs
Enable for all regions: Yes
Log file validation: Enable (detects if log files are tampered with)

Combine with Athena to query months of API history with SQL. Who deleted that Security Group? Who changed that IAM policy at 3 AM last Tuesday? CloudTrail has the answer. Athena makes it searchable in seconds.

AWS Config — Compliance Recording

Config records the configuration of your AWS resources over time and checks whether they comply with rules you define. It does not prevent changes — it detects and reports them.

What Config answers:

TEXT
Is there unrestricted SSH access on any Security Group right now?
Do all EBS volumes have encryption enabled?
Which EC2 instances are missing required tags?
How has this ALB configuration changed over the last 30 days?
Was this S3 bucket compliant last Tuesday?

Config Rules — define what compliant looks like:

AWS provides 75+ managed rules. You can also write custom rules using Lambda.

◈ DIAGRAM
ec2-instance-no-public-ip → flag EC2 with public IPs
restricted-ssh → flag Security Groups allowing 0.0.0.0/0 on port 22
s3-bucket-versioning-enabled → flag S3 buckets without versioning
root-account-mfa-enabled → flag accounts with no MFA on root
encrypted-volumes → flag unencrypted EBS volumes
required-tags → flag resources missing required tags

Rules evaluate on every configuration change and at regular time intervals.

Remember

Config Rules do not prevent actions. If someone opens port 22 to 0.0.0.0/0, Config detects and flags it — but the change already happened. For prevention use IAM Deny policies or SCPs. Use Config for continuous detection and automated remediation after the fact.

Automatic Remediation:

When a rule flags a non-compliant resource, Config can automatically fix it using SSM Automation Documents.

◈ DIAGRAM
Security Group opens port 22 to 0.0.0.0/0 → NON_COMPLIANT
↓
Config detects via rule
↓
SSM Automation runs: removes the offending inbound rule
↓
Resource returns to compliant state
Notification sent to ops team

CloudWatch vs CloudTrail vs Config — Which Does What

These three services are constantly confused because they all deal with visibility. They serve completely different purposes.

CloudWatch CloudTrail Config
Question it answers Is my system healthy right now? Who did what and when? Are my resources configured correctly?
Watches Metrics, logs, performance data Every API call in your account Resource configuration state
Data type Time-series numbers and log text Structured API call records Configuration snapshots and change history
Alerts via Alarms → SNS, ASG, EC2 recovery EventBridge → SNS, Lambda EventBridge → SNS

Real example using one Load Balancer:

◈ DIAGRAM
CloudWatch → monitoring incoming request count, 5xx error rate, latency
Alarm fires when error rate exceeds 5%
Config → tracking Security Group changes on the ALB
Flags if SSL certificate expires
Records every listener rule change over 90 days
CloudTrail → recording exactly who changed the listener routing rules
At what time, from which IP, via which tool

All three run simultaneously. All three answer different questions. None of them replaces the others.

Hands-on Lab — CloudWatch Alarm, CloudTrail Trail, Config Rule

Step 1 — Create SNS topic for alerts

◈ DIAGRAM
SNS → Topics → Create topic
Type: Standard
Name: devops-monitoring-alerts → Create topic
Subscribe your email:
Subscriptions → Create subscription
Protocol: Email → Endpoint: your email → Create subscription
Confirm the subscription from your inbox before continuing

Step 2 — Create a CloudWatch Alarm

◈ DIAGRAM
CloudWatch → Alarms → Create alarm → Select metric
Browse: EC2 → Per-Instance Metrics → CPUUtilization → select an instance → Select metric
Conditions:
Threshold type: Static
Whenever CPUUtilization is: Greater than 80
Period: 5 minutes
Actions:
Notification → In alarm → select devops-monitoring-alerts SNS topic
Alarm name: high-cpu-prod
Create alarm

Step 3 — Test the alarm immediately

Bash
aws cloudwatch set-alarm-state \
--alarm-name "high-cpu-prod" \
--state-value ALARM \
--state-reason "Testing alarm pipeline" \
--region ap-south-1
◈ DIAGRAM
Check your email — notification should arrive within 60 seconds.
CloudWatch → Alarms → high-cpu-prod → State should show: In alarm

Step 4 — Enable CloudTrail

◈ DIAGRAM
CloudTrail → Create trail
Trail name: devops-audit-trail
Storage location: Create new S3 bucket → devops-cloudtrail-ACCOUNTID
Log file SSE-KMS encryption: optional for lab
Enable for all regions: Yes → tick
Log file validation: Enable → Create trail
CloudTrail → Trails → devops-audit-trail → Logging: On

Step 5 — Enable Data Events for an S3 bucket

◈ DIAGRAM
CloudTrail → devops-audit-trail → Edit
Events: Additional settings
Data events → Add data event type → S3
Select your S3 bucket → Log Read and Write events → Save

Step 6 — Enable AWS Config

◈ DIAGRAM
Config → Get started (or Settings if already set up)
Recording method: Record all resources
S3 bucket: Create new → devops-config-ACCOUNTID
Service role: Create Config service-linked role → Next → Confirm
Config → Rules → Add rule
Search: restricted-ssh → Add rule → Save
Config → Rules → restricted-ssh
Check: Compliance status
Any Security Group with 0.0.0.0/0 on port 22 will show as NON_COMPLIANT

Step 7 — Verify CloudTrail is recording

◈ DIAGRAM
Perform any action — create an S3 bucket, modify a Security Group
CloudTrail → Event history → filter by Event name
Search for: CreateBucket or AuthorizeSecurityGroupIngress
You should see the event with: user name, source IP, time

Step 8 — Cleanup

◈ DIAGRAM
CloudWatch → Alarms → high-cpu-prod → Delete alarm
CloudTrail → Trails → devops-audit-trail → Stop logging → Delete trail
Config → Settings → Stop recording (do not delete rules yet)
S3 → delete cloudtrail and config buckets (must empty first)
SNS → Topics → devops-monitoring-alerts → Delete

Common Mistakes to Avoid

Common Mistake

Confusing CloudWatch with CloudTrail. CloudWatch answers "is my system healthy right now?" — it watches metrics and performance. CloudTrail answers "who made this API call and when?" — it records every action. Both should run simultaneously. They are not alternatives, they are complementary.

Common Mistake

Assuming CloudTrail logs everything by default. Management events yes — creating EC2, modifying IAM, changing Security Groups. Data events no — S3 object reads and writes, Lambda invocations are NOT logged unless you explicitly enable them. For a compliance-critical S3 bucket, this is a critical gap.

Common Mistake

Treating Config Rules as preventive controls. Config detects non-compliance after it happens — it does not block the action. Opening port 22 to 0.0.0.0/0? Config detects and flags it. But the port is already open. Use IAM Deny policies or SCPs for prevention. Use Config for detection and automated remediation.

Security

Store CloudTrail logs in a separate dedicated S3 bucket with Object Lock enabled. If an attacker compromises your main account, they cannot delete CloudTrail logs locked in a separate account. This is the only way to guarantee audit log integrity after a breach.

Tip

The fastest way to answer "who changed this?" in AWS is CloudTrail + Athena. Create an Athena table over your CloudTrail S3 bucket and query months of API call history with SQL. Who deleted that Security Group rule? Who modified that IAM policy at 3 AM? CloudTrail has it. Athena surfaces it in seconds.

Resources

AWS Direct Connect vs Site-to-Site VPN Failover

AWS Direct Connect vs Site-to-Site VPN Failover

Direct Connect vs VPN isn't really either/or for production — it's a primary-plus-failover pattern. Here's how to design it, and when either/or is right.

5 min read•Aug 2026
Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot: The Cost Crossover

Lambda vs Fargate vs EC2 Spot, at the crossover where Lambda stops being cheaper — 2026 pricing, invocation thresholds, and interruption math.

5 min read•Aug 2026
Secrets Manager vs Parameter Store vs Vault

Secrets Manager vs Parameter Store vs Vault

AWS Secrets Manager, Parameter Store, and HashiCorp Vault compared for 2026 - cost math, rotation, multi-cloud fit, and the Vault-to-OpenBao fork.

5 min read•Aug 2026
AWS VPC Security: Hardening Every Layer

AWS VPC Security: Hardening Every Layer

Most cloud security incidents start with a misconfigured VPC. Here's how to harden every layer — subnets, Security Groups, NACLs, and IAM — for production.

5 min read•Jul 2026
Event-Driven Architecture on AWS Explained

Event-Driven Architecture on AWS Explained

Event-driven architecture on AWS decouples services and absorbs traffic spikes using SQS, SNS, EventBridge, and Lambda — workflows that scale themselves.

5 min read•Jul 2026
S3 vs RDS vs DynamoDB: Choosing AWS Storage

S3 vs RDS vs DynamoDB: Choosing AWS Storage

Choosing S3, RDS, or DynamoDB wrong costs you in performance, cost, and scalability. Here is a practical decision guide based on your actual access patterns.

5 min read•Jul 2026
AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS Cost Optimisation: Cut Cloud Bills 40-60%

AWS bills surprise teams every month. Here are the 8 concrete actions that cut cloud spend by 40-60% without touching your application architecture.

5 min read•Jul 2026
EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2 vs Lambda vs Fargate: Choosing AWS Compute

EC2, Lambda, or Fargate — choosing the wrong AWS compute option costs you money and performance. Here is exactly when to use each one in production.

5 min read•Jul 2026

Explore More in AWS DevOps, Cost, and Machine Learning

All 6 Topics

Frequently Asked Questions

Is CloudWatch, CloudTrail, and Config - Monitoring, Auditing, and Compliance free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the CloudWatch, CloudTrail, and Config - Monitoring, Auditing, and Compliance topic cover?

Monitor AWS infrastructure with CloudWatch metrics and alarms, audit every API call with CloudTrail, and enforce compliance rules with AWS Config.