Skip to main content

OpenTelemetry and Distributed Tracing

Learn to instrument services with OpenTelemetry, run the Collector, and follow one slow request across services to find the real cause of an incident.

~3 hours
11 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Metrics and Logs Leave a Blind Spot

You are the platform engineer at acme-shop. In the GitOps module you deployed cart-service from Git and shipped a canary.

Understanding Traces, Spans, and Context Propagation

What a trace and a span are A trace is the record of one request from the moment it arrives until the response leaves.

Instrumenting Python Services with OpenTelemetry

Installing the packages The Python SDK is split into an API and an SDK. Your code imports from the API. The SDK is configured once at startup.

Using Semantic Conventions for Consistent Attributes

Why standard attribute names matter Semantic conventions are OpenTelemetry's shared dictionary of attribute names.

Running the OpenTelemetry Collector

What the Collector does and why you need it The OpenTelemetry Collector is a standalone service between your apps and your backends.

Choosing a Tracing Backend

Jaeger for learning and small setups Jaeger is an open source tracing backend and UI.

Skills You'll Master

OPENTELEMETRYDISTRIBUTED-TRACINGOBSERVABILITYOTEL-COLLECTORJAEGER

Curriculum Index11 topics

1

Understanding Why Metrics and Logs Leave a Blind Spot

You are the platform engineer at acme-shop.

2

Understanding Traces, Spans, and Context Propagation

What a trace and a span are A trace is the record of one request from the moment it arrives until the response leaves.

3

Instrumenting Python Services with OpenTelemetry

Installing the packages The Python SDK is split into an API and an SDK. Your code imports from the API.

4

Using Semantic Conventions for Consistent Attributes

Why standard attribute names matter Semantic conventions are OpenTelemetry's shared dictionary of attribute names.

5

Running the OpenTelemetry Collector

What the Collector does and why you need it The OpenTelemetry Collector is a standalone service between your apps and...

6

Choosing a Tracing Backend

Jaeger for learning and small setups Jaeger is an open source tracing backend and UI.

7

Sampling Traces Without Losing the Interesting Ones

Why you cannot keep everything A service handling 10,000 requests per minute with five spans each produces 50,000 spans...

8

Connecting Traces to Logs and Metrics

Why correlation matters A trace tells you one request was slow.

9

Debugging an Incident With Traces

Reading a trace waterfall A trace backend draws the waterfall: one bar per span, indented by parent and child.

10

Tracing a Slow Checkout in a Hands-On Lab

📌 Remember: This lab runs on your laptop with Docker Compose, so it costs nothing.

11

Quick Reference and Common Mistakes

Quick reference Common mistakes Mounting the Collector config at the wrong path is easy to do because the container...

Career Impact

Roles that use the skills in this module.

  • SRE

    ₹20L - ₹38L a year

    High Demand
  • AIOps Engineer

    ₹18L - ₹35L a year

    High Demand
  • Platform Engineer

    ₹15L - ₹28L a year

    Growing
See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

Distributed tracing records the path of one request as it moves through several services, with the timing of each step. A tracing backend joins the steps using a shared trace ID, so you can see which call was slow or failed.

You can export straight from an app to a backend, but the Collector is worth it once you have more than a couple of services. It gives every service one place to send data, and you change backends in one config file instead of in every service.

Context propagation broke somewhere. The caller did not send the traceparent header, or the receiver did not read it, so the second service started a new trace. Check the HTTP client, proxies, and any queue between the services.

Usually yes, once traffic is high. Keep a small share of normal traces and use tail sampling to keep every error and slow trace. Logs are not sampled, so you still keep all of them.