Monitoring and Logging: Prometheus to Tracing
Learn DevOps monitoring and logging from zero: Prometheus, PromQL, Grafana, alerting, Loki, EFK, tracing, Kubernetes and SLOs, with a full hands-on lab.
What You'll Learn
Understanding Why Monitoring and Logging Matter
9:40 PM on the first night of a big sale. Orders are pouring in, and then quietly, about one in five checkouts starts failing.
Understanding Observability Fundamentals
What should you measure first? There are thousands of possible metrics, so start with the ones that reflect what users feel.
Understanding Prometheus Architecture
What is Prometheus and how does it collect data?
Collecting Metrics with Exporters and Instrumentation
What are exporters and which ones matter? Most software does not speak Prometheus natively.
Querying Metrics with PromQL
How do you select metrics? PromQL (Prometheus Query Language) is how you ask questions about metrics, in the Prometheus UI, in Grafana panels and in...
Visualizing Metrics with Grafana
What is Grafana and why use it? Grafana is an open-source tool for building dashboards and exploring data.
Skills You'll Master
Curriculum Index14 topics
Understanding Why Monitoring and Logging Matter
9:40 PM on the first night of a big sale.
Understanding Observability Fundamentals
What should you measure first? There are thousands of possible metrics, so start with the ones that reflect what users...
Understanding Prometheus Architecture
What is Prometheus and how does it collect data?
Collecting Metrics with Exporters and Instrumentation
What are exporters and which ones matter? Most software does not speak Prometheus natively.
Querying Metrics with PromQL
How do you select metrics? PromQL (Prometheus Query Language) is how you ask questions about metrics, in the Prometheus...
Visualizing Metrics with Grafana
What is Grafana and why use it? Grafana is an open-source tool for building dashboards and exploring data.
Alerting with Prometheus, Alertmanager and Grafana
How do Prometheus alert rules work? An alert rule is a PromQL expression plus a duration.
Centralizing Logs with Loki and Grafana Alloy
Why centralize logs, and what makes a good log line?
Searching Logs with the EFK and ELK Stack
What are the ELK and EFK stacks? The ELK stack is three tools from Elastic: Elasticsearch stores and indexes logs...
Tracing Requests with OpenTelemetry and Jaeger
What is distributed tracing? In a microservices system, one click can pass through ten services.
Monitoring Kubernetes with kube-prometheus-stack
How do you install monitoring on a Kubernetes cluster?
Measuring Reliability with SLIs, SLOs and Error Budgets
What are SLIs, SLOs and SLAs? Without a definition of "healthy", every incident becomes a debate.
Building an Observability Stack Hands-On
What will you build in this lab? You will run a complete observability stack with Docker Compose: a small Python shop...
Securing, Troubleshooting and Scaling Your Monitoring Stack
How do you secure the monitoring stack? Monitoring data is sensitive.
Career Impact
Roles that use the skills in this module.
Site Reliability Engineer
DevOps Engineer
Platform Engineer
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
Monitoring tracks numbers over time, such as request rate, error rate and CPU, and tells you that something is wrong. Logging records individual events with details, and tells you why it went wrong. You need both: metrics detect problems and trigger alerts, logs explain them.
Start with Prometheus and Grafana. They are free, run anywhere, are the default in Kubernetes, and teach the concepts every other tool uses. Once you understand metrics, PromQL and alerting, moving to Datadog or New Relic is mostly learning a new interface.
Observability is how well you can understand what is happening inside a system from the data it produces. It is built on three kinds of telemetry: metrics, logs and traces. Monitoring answers known questions; observability helps you investigate problems you did not predict.
Neither is better for everything. Loki indexes only labels, so it is cheaper and simpler, and fits naturally with Prometheus and Grafana. Elasticsearch indexes every word, so it is stronger for full-text search and analytics but needs much more memory and care to run.