What you will learn
- CloudWatch Metrics — namespaces, dimensions, and why RAM is not collected by default
- Custom Metrics — pushing your own business data into CloudWatch
- CloudWatch Logs — log groups, streams, and the Unified Agent
- Subscription Filters — shipping logs in real time to Lambda, Kinesis, or Firehose
- CloudWatch Logs Insights — querying your logs like a database
- CloudWatch Alarms — states, composite alarms, and EC2 auto-recovery
- The four CloudWatch Insights features — Container, Lambda, Contributor, Application
- Amazon EventBridge — scheduling tasks and reacting to AWS events automatically
- AWS CloudTrail — who did what, when, and from where
- AWS Config — are my resources configured correctly right now
- CloudWatch vs CloudTrail vs Config — which service answers which question
Why this matters
A Zerodha trading platform going down at market open. A Razorpay payment API throwing 500 errors. An EC2 instance being terminated by an unknown IAM user at 2 AM. None of these can be investigated without the right visibility tools in place.
CloudWatch tells you something is wrong right now. CloudTrail tells you who did what and when. Config tells you whether your infrastructure drifted from its intended state. Together they form the observability and compliance backbone of every serious AWS deployment. Without them you are flying blind — debugging by guesswork instead of data.
CloudWatch Metrics — What AWS Measures Automatically
CloudWatch collects metrics from every AWS service automatically. A metric is a numeric value tracked over time — CPU utilisation, network bytes, request count, disk reads.
Key concepts:
Namespace → the folder that groups metrics per service EC2 metrics live in: AWS/EC2 Lambda metrics live in: AWS/Lambda RDS metrics live in: AWS/RDS Dimension → extra context that identifies which specific resource InstanceId = i-0abc123 (which EC2) FunctionName = devops-orders (which Lambda) DBInstanceIdentifier = prod-db (which RDS) Period → how often the metric is evaluated Standard: every 60 seconds High resolution: every 1, 5, 10, or 30 secondsRememberRAM (memory usage) is NOT collected by EC2 by default. CloudWatch gets CPU, network, disk I/O, and status checks from EC2 automatically — but memory requires the CloudWatch Unified Agent installed on the instance. This is one of the most commonly tested CloudWatch facts.
Custom Metrics — push your own data:
Any number you care about can become a CloudWatch metric. Active users, queue depth, business transaction count, cache hit rate — anything your application can measure.
## Push a custom metric — active sessions countaws cloudwatch put-metric-data \ --namespace "DevOpsNetwork/App" \ --metric-name "ActiveSessions" \ --value 1247 \ --unit Count \ --dimensions Environment=prod,Service=orders \ --region ap-south-1Custom metrics appear in CloudWatch alongside AWS metrics and can trigger alarms just like any built-in metric.
CloudWatch Logs — Centralised Log Storage
CloudWatch Logs stores application logs, infrastructure logs, and audit logs from across your entire AWS environment.
Core structure:
Log Group → one per application or service /aws/lambda/devops-orders /aws/rds/instance/prod-db/error /vpc/prod-vpc-flow-logs Log Stream → logs from one specific source inside the group A new stream per Lambda execution environment One stream per EC2 instance Retention → how long logs are kept Default: never expire (costs money over time) Set explicitly: 1 day to 10 years, or never expire Best practice: set retention on every log groupWhat sends logs to CloudWatch automatically:
Lambda → all print() and console.log() outputAPI Gateway → access logs (when enabled)VPC Flow Logs → network traffic recordsCloudTrail → API call audit eventsRDS → error logs, slow query logs (when enabled)EC2 sends NO logs by default. You must install the CloudWatch Agent.
Logs Agent vs Unified Agent:
| Logs Agent | Unified Agent | |
|---|---|---|
| Version | Old | New — always use this |
| Sends logs | Yes | Yes |
| Sends system metrics | No | Yes — RAM, CPU detail, disk, netstat |
| Config management | Manual config file | Via SSM Parameter Store (centralised) |
Always use the Unified Agent. It does everything the old agent does plus gives you RAM usage, process counts, disk usage, TCP connections — metrics EC2 does not expose natively.
Where logs can go from CloudWatch:
CloudWatch Logs ├── S3 (archiving — up to 12 hour delay, use for long-term storage) ├── Kinesis Data Streams (real-time to your own consumers) ├── Kinesis Data Firehose (near real-time to S3, Redshift, OpenSearch) └── Lambda (trigger a function when a log pattern matches)RememberExporting to S3 (CreateExportTask) takes up to 12 hours. It is for archiving, not real-time alerting. For real-time log processing use Subscription Filters — they deliver instantly.
Subscription Filters — Real-Time Log Routing
A Subscription Filter watches a log group and forwards matching log events in real time to a destination.
/aws/lambda/devops-orders → filter pattern: "ERROR" ↓Only ERROR log lines forwarded → Lambda function ↓Lambda sends Slack alert immediatelyMulti-account log aggregation:
Production, staging, and dev accounts each forward logs via Subscription Filters to a central Kinesis Data Streams in a security account → Kinesis Firehose → centralised S3 bucket. One bucket, all logs, searchable with Athena.
CloudWatch Logs Insights — Query Your Logs
Logs Insights lets you run SQL-like queries against your stored log data. Filter, aggregate, sort, count — across one or multiple log groups.
## Find all errors in the last hourfields @timestamp, @message| filter @message like /ERROR/| sort @timestamp desc| limit 50 ## Count errors per 5-minute windowfields @timestamp, @message| filter @message like /Exception/| stats count(*) as errors by bin(5m) ## Find slowest Lambda invocationsfilter @type = "REPORT"| stats avg(@duration), max(@duration) by bin(5m)| sort max(@duration) descResults can be saved as named queries and pinned directly to a CloudWatch Dashboard.
RememberLogs Insights queries historical data already in CloudWatch. It is not real-time streaming. For real-time log processing use Subscription Filters + Lambda.
CloudWatch Alarms — Act When Something Goes Wrong
An alarm watches a single metric and fires when it crosses a threshold you define.
Three states:
OK → metric is within acceptable rangeALARM → threshold breachedINSUFFICIENT_DATA → not enough data points yet to evaluateWhat an alarm can trigger:
EC2 Action → stop, terminate, reboot, or recover the instanceAuto Scaling → scale out or scale inSNS → send notification (email, Slack via Lambda, PagerDuty)## Create alarm: fire when average CPU > 80% for 5 minutesaws cloudwatch put-metric-alarm \ --alarm-name "high-cpu-prod" \ --metric-name CPUUtilization \ --namespace AWS/EC2 \ --dimensions Name=InstanceId,Value=i-0abc123 \ --statistic Average \ --period 300 \ --threshold 80 \ --comparison-operator GreaterThanThreshold \ --evaluation-periods 1 \ --alarm-actions arn:aws:sns:ap-south-1:123456789012:devops-alerts \ --region ap-south-1 ## Test immediately without waiting for real CPU spikeaws cloudwatch set-alarm-state \ --alarm-name "high-cpu-prod" \ --state-value ALARM \ --state-reason "testing" \ --region ap-south-1Composite Alarms — reduce alert fatigue:
A single alarm firing on every small CPU spike creates noise. Engineers start ignoring pages. A Composite Alarm combines multiple alarms with AND/OR logic.
Alarm A (CPU > 80%) ─┐ ├──→ Composite Alarm (CPU AND Memory both high) ──→ SNS pageAlarm B (Memory > 85%) ─┘ CPU spikes briefly alone → Composite stays OK → no pageCPU AND Memory both high → Composite fires → page the on-call engineerComposite Alarms dramatically reduce alert noise. You only get paged when a real problem is confirmed.
EC2 Instance Recovery:
CloudWatch can automatically recover a broken EC2 instance when the underlying hardware fails.
CloudWatch monitors EC2 status checks ↓StatusCheckFailed_System alarm fires (hardware issue, not app issue) ↓EC2 Instance Recovery triggered automatically ↓Instance migrated to healthy hardwareKeeps: same private IP, same public IP, same Elastic IP, same metadataNo data loss. No human action. The instance recovers itself.
CloudWatch Insights — Four Specialised Features
Container Insights:
Collects metrics and logs from ECS, EKS, and Kubernetes. Shows cluster, node, pod, and task-level visibility in pre-built dashboards. For EKS, CloudWatch deploys a containerised agent inside the cluster automatically.
Lambda Insights:
Monitoring specifically for Lambda. Collects CPU time, memory, disk, network, cold start count, and worker shutdowns. Deployed as a Lambda Layer — attach to your function, no code changes needed.
Contributor Insights:
Reads log data and finds the top contributors to a metric. Who or what is causing the most traffic, errors, or load?
Example: find the top 10 IPs generating the most 404 errors in ALB access logsExample: find the top 5 DynamoDB partition keys being accessed most frequentlyApplication Insights:
Automated monitoring for your entire application stack. Point it at your EC2 application resources, it discovers related components (RDS, ELB, Auto Scaling), creates dashboards automatically, and sends findings when problems occur. Uses machine learning for anomaly detection.
Amazon EventBridge — React to Everything
EventBridge is an event router. Something happens anywhere in AWS — EC2 state changes, S3 uploads, CodeBuild failures, IAM actions — EventBridge sees it and routes it to a target.
Two ways to use it:
Schedule → run something on a cron or rate-based timer Every day at 2 AM → EventBridge → Lambda → delete old logs Every 5 minutes → EventBridge → Lambda → check health of external API Event Pattern → react when something specific happens Root user signs in → EventBridge → SNS → email alert immediately EC2 instance stopped → EventBridge → Lambda → send Slack notification S3 object deleted → EventBridge → Lambda → log the deletion for auditEvent buses:
Default Event Bus → all AWS service events (EC2, S3, Lambda, etc)Partner Event Bus → third-party SaaS events (Datadog, Zendesk, Salesforce)Custom Event Bus → your own application eventsSchema Registry:
EventBridge automatically detects the structure of events flowing through it. Your application can download a schema and generate typed code that already knows the field names and types — no manual inspection needed.
Real production example:
IAM root user logs in at 2 AM ↓CloudTrail records the event ↓EventBridge rule matches: source=aws.signin, userIdentity.type=Root ↓SNS sends email to security team immediately ↓Lambda automatically disables the root access key (if one exists)All of this happens automatically. Nobody manually watched CloudTrail.
AWS CloudTrail — The Audit Log
CloudTrail records every API call made in your AWS account — console clicks, CLI commands, SDK calls, and service-to-service calls. It answers: who did what, when, and from where.
CloudTrail is enabled by defaultEvents stored free for 90 days in CloudTrail Event HistoryFor longer retention: create a Trail to send events to S3Three event types:
Management Events (logged by default): Operations performed on AWS resources Examples: creating an EC2, modifying a Security Group, adding an IAM policy Read and Write events can be separated to reduce noise Data Events (NOT logged by default — high volume): Object-level operations inside resources Examples: S3 GetObject, PutObject, DeleteObject; Lambda Invoke Must be explicitly enabled — extra cost Critical for: compliance on sensitive buckets, auditing Lambda invocations CloudTrail Insights Events: Unusual activity detection using baseline analysis Detects: unexpected API call spikes, unusual IAM action bursts When anomaly found → stored in S3, visible in console, EventBridge event firedRememberS3 object-level actions (GetObject, PutObject, DeleteObject) are Data Events and are NOT logged by default. Teams regularly miss this when setting up compliance logging. Always explicitly enable Data Events for any S3 bucket containing sensitive data.
Storing trails long-term:
CloudTrail → Trails → Create TrailTrail name: prod-audit-trailS3 bucket: create new → devops-cloudtrail-logsEnable for all regions: YesLog file validation: Enable (detects if log files are tampered with)Combine with Athena to query months of API history with SQL. Who deleted that Security Group? Who changed that IAM policy at 3 AM last Tuesday? CloudTrail has the answer. Athena makes it searchable in seconds.
AWS Config — Compliance Recording
Config records the configuration of your AWS resources over time and checks whether they comply with rules you define. It does not prevent changes — it detects and reports them.
What Config answers:
Is there unrestricted SSH access on any Security Group right now?Do all EBS volumes have encryption enabled?Which EC2 instances are missing required tags?How has this ALB configuration changed over the last 30 days?Was this S3 bucket compliant last Tuesday?Config Rules — define what compliant looks like:
AWS provides 75+ managed rules. You can also write custom rules using Lambda.
ec2-instance-no-public-ip → flag EC2 with public IPsrestricted-ssh → flag Security Groups allowing 0.0.0.0/0 on port 22s3-bucket-versioning-enabled → flag S3 buckets without versioningroot-account-mfa-enabled → flag accounts with no MFA on rootencrypted-volumes → flag unencrypted EBS volumesrequired-tags → flag resources missing required tagsRules evaluate on every configuration change and at regular time intervals.
RememberConfig Rules do not prevent actions. If someone opens port 22 to 0.0.0.0/0, Config detects and flags it — but the change already happened. For prevention use IAM Deny policies or SCPs. Use Config for continuous detection and automated remediation after the fact.
Automatic Remediation:
When a rule flags a non-compliant resource, Config can automatically fix it using SSM Automation Documents.
Security Group opens port 22 to 0.0.0.0/0 → NON_COMPLIANT ↓Config detects via rule ↓SSM Automation runs: removes the offending inbound rule ↓Resource returns to compliant stateNotification sent to ops teamCloudWatch vs CloudTrail vs Config — Which Does What
These three services are constantly confused because they all deal with visibility. They serve completely different purposes.
| CloudWatch | CloudTrail | Config | |
|---|---|---|---|
| Question it answers | Is my system healthy right now? | Who did what and when? | Are my resources configured correctly? |
| Watches | Metrics, logs, performance data | Every API call in your account | Resource configuration state |
| Data type | Time-series numbers and log text | Structured API call records | Configuration snapshots and change history |
| Alerts via | Alarms → SNS, ASG, EC2 recovery | EventBridge → SNS, Lambda | EventBridge → SNS |
Real example using one Load Balancer:
CloudWatch → monitoring incoming request count, 5xx error rate, latency Alarm fires when error rate exceeds 5% Config → tracking Security Group changes on the ALB Flags if SSL certificate expires Records every listener rule change over 90 days CloudTrail → recording exactly who changed the listener routing rules At what time, from which IP, via which toolAll three run simultaneously. All three answer different questions. None of them replaces the others.
Hands-on Lab — CloudWatch Alarm, CloudTrail Trail, Config Rule
Step 1 — Create SNS topic for alerts
SNS → Topics → Create topicType: StandardName: devops-monitoring-alerts → Create topic Subscribe your email:Subscriptions → Create subscriptionProtocol: Email → Endpoint: your email → Create subscriptionConfirm the subscription from your inbox before continuingStep 2 — Create a CloudWatch Alarm
CloudWatch → Alarms → Create alarm → Select metricBrowse: EC2 → Per-Instance Metrics → CPUUtilization → select an instance → Select metric Conditions:Threshold type: StaticWhenever CPUUtilization is: Greater than 80Period: 5 minutes Actions:Notification → In alarm → select devops-monitoring-alerts SNS topic Alarm name: high-cpu-prodCreate alarmStep 3 — Test the alarm immediately
aws cloudwatch set-alarm-state \ --alarm-name "high-cpu-prod" \ --state-value ALARM \ --state-reason "Testing alarm pipeline" \ --region ap-south-1Check your email — notification should arrive within 60 seconds.CloudWatch → Alarms → high-cpu-prod → State should show: In alarmStep 4 — Enable CloudTrail
CloudTrail → Create trailTrail name: devops-audit-trailStorage location: Create new S3 bucket → devops-cloudtrail-ACCOUNTIDLog file SSE-KMS encryption: optional for labEnable for all regions: Yes → tickLog file validation: Enable → Create trail CloudTrail → Trails → devops-audit-trail → Logging: OnStep 5 — Enable Data Events for an S3 bucket
CloudTrail → devops-audit-trail → EditEvents: Additional settingsData events → Add data event type → S3Select your S3 bucket → Log Read and Write events → SaveStep 6 — Enable AWS Config
Config → Get started (or Settings if already set up)Recording method: Record all resourcesS3 bucket: Create new → devops-config-ACCOUNTIDService role: Create Config service-linked role → Next → Confirm Config → Rules → Add ruleSearch: restricted-ssh → Add rule → Save Config → Rules → restricted-sshCheck: Compliance statusAny Security Group with 0.0.0.0/0 on port 22 will show as NON_COMPLIANTStep 7 — Verify CloudTrail is recording
Perform any action — create an S3 bucket, modify a Security Group CloudTrail → Event history → filter by Event nameSearch for: CreateBucket or AuthorizeSecurityGroupIngressYou should see the event with: user name, source IP, timeStep 8 — Cleanup
CloudWatch → Alarms → high-cpu-prod → Delete alarmCloudTrail → Trails → devops-audit-trail → Stop logging → Delete trailConfig → Settings → Stop recording (do not delete rules yet)S3 → delete cloudtrail and config buckets (must empty first)SNS → Topics → devops-monitoring-alerts → DeleteCommon Mistakes to Avoid
Common MistakeConfusing CloudWatch with CloudTrail. CloudWatch answers "is my system healthy right now?" — it watches metrics and performance. CloudTrail answers "who made this API call and when?" — it records every action. Both should run simultaneously. They are not alternatives, they are complementary.
Common MistakeAssuming CloudTrail logs everything by default. Management events yes — creating EC2, modifying IAM, changing Security Groups. Data events no — S3 object reads and writes, Lambda invocations are NOT logged unless you explicitly enable them. For a compliance-critical S3 bucket, this is a critical gap.
Common MistakeTreating Config Rules as preventive controls. Config detects non-compliance after it happens — it does not block the action. Opening port 22 to 0.0.0.0/0? Config detects and flags it. But the port is already open. Use IAM Deny policies or SCPs for prevention. Use Config for detection and automated remediation.
SecurityStore CloudTrail logs in a separate dedicated S3 bucket with Object Lock enabled. If an attacker compromises your main account, they cannot delete CloudTrail logs locked in a separate account. This is the only way to guarantee audit log integrity after a breach.
TipThe fastest way to answer "who changed this?" in AWS is CloudTrail + Athena. Create an Athena table over your CloudTrail S3 bucket and query months of API call history with SQL. Who deleted that Security Group rule? Who modified that IAM policy at 3 AM? CloudTrail has it. Athena surfaces it in seconds.