Skip to main content

Monitoring and Logging: Prometheus to Tracing

Learn DevOps monitoring and logging from zero: Prometheus, PromQL, Grafana, alerting, Loki, EFK, tracing, Kubernetes and SLOs, with a full hands-on lab.

~3 hours
14 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Monitoring and Logging Matter

9:40 PM on the first night of a big sale. Orders are pouring in, and then quietly, about one in five checkouts starts failing.

Understanding Observability Fundamentals

What should you measure first? There are thousands of possible metrics, so start with the ones that reflect what users feel.

Understanding Prometheus Architecture

What is Prometheus and how does it collect data?

Collecting Metrics with Exporters and Instrumentation

What are exporters and which ones matter? Most software does not speak Prometheus natively.

Querying Metrics with PromQL

How do you select metrics? PromQL (Prometheus Query Language) is how you ask questions about metrics, in the Prometheus UI, in Grafana panels and in...

Visualizing Metrics with Grafana

What is Grafana and why use it? Grafana is an open-source tool for building dashboards and exploring data.

Skills You'll Master

PROMETHEUSGRAFANALOGGINGOBSERVABILITYALERTINGOPENTELEMETRY

Curriculum Index14 topics

1

Understanding Why Monitoring and Logging Matter

9:40 PM on the first night of a big sale.

2

Understanding Observability Fundamentals

What should you measure first? There are thousands of possible metrics, so start with the ones that reflect what users...

3

Understanding Prometheus Architecture

What is Prometheus and how does it collect data?

4

Collecting Metrics with Exporters and Instrumentation

What are exporters and which ones matter? Most software does not speak Prometheus natively.

5

Querying Metrics with PromQL

How do you select metrics? PromQL (Prometheus Query Language) is how you ask questions about metrics, in the Prometheus...

6

Visualizing Metrics with Grafana

What is Grafana and why use it? Grafana is an open-source tool for building dashboards and exploring data.

7

Alerting with Prometheus, Alertmanager and Grafana

How do Prometheus alert rules work? An alert rule is a PromQL expression plus a duration.

8

Centralizing Logs with Loki and Grafana Alloy

Why centralize logs, and what makes a good log line?

9

Searching Logs with the EFK and ELK Stack

What are the ELK and EFK stacks? The ELK stack is three tools from Elastic: Elasticsearch stores and indexes logs...

10

Tracing Requests with OpenTelemetry and Jaeger

What is distributed tracing? In a microservices system, one click can pass through ten services.

11

Monitoring Kubernetes with kube-prometheus-stack

How do you install monitoring on a Kubernetes cluster?

12

Measuring Reliability with SLIs, SLOs and Error Budgets

What are SLIs, SLOs and SLAs? Without a definition of "healthy", every incident becomes a debate.

13

Building an Observability Stack Hands-On

What will you build in this lab? You will run a complete observability stack with Docker Compose: a small Python shop...

14

Securing, Troubleshooting and Scaling Your Monitoring Stack

How do you secure the monitoring stack? Monitoring data is sensitive.

Career Impact

Roles that use the skills in this module.

  • Site Reliability Engineer

  • DevOps Engineer

  • Platform Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

Monitoring tracks numbers over time, such as request rate, error rate and CPU, and tells you that something is wrong. Logging records individual events with details, and tells you why it went wrong. You need both: metrics detect problems and trigger alerts, logs explain them.

Start with Prometheus and Grafana. They are free, run anywhere, are the default in Kubernetes, and teach the concepts every other tool uses. Once you understand metrics, PromQL and alerting, moving to Datadog or New Relic is mostly learning a new interface.

Observability is how well you can understand what is happening inside a system from the data it produces. It is built on three kinds of telemetry: metrics, logs and traces. Monitoring answers known questions; observability helps you investigate problems you did not predict.

Neither is better for everything. Loki indexes only labels, so it is cheaper and simpler, and fits naturally with Prometheus and Grafana. Elasticsearch indexes every word, so it is stronger for full-text search and analytics but needs much more memory and care to run.