Skip to main content

SLOs and Observability for SRE

Learn to define SLIs and SLOs, track error budgets, and write multi-window burn-rate alerts that page only when users are really affected.

~3.5 hours
12 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Dashboards Are Not Enough

It is the Friday of acme-shop's biggest sale. Every Grafana panel is green: CPU at 40 percent, every pod Running, no restarts.

Choosing SLIs That Reflect User Pain

What an SLI is and the three common shapes An SLI (Service Level Indicator) is one precisely defined measurement of something users experience...

Setting SLOs That Match What Users Need

What an SLO is An SLO (Service Level Objective) is the target value for an SLI over a defined time window.

Spending and Protecting the Error Budget

What an error budget is The error budget is the unreliability your SLO allows: one minus the SLO, applied to the events in the window.

Understanding Burn Rate and Why One Threshold Fails

What burn rate means Burn rate is how fast you are spending the error budget compared with the pace that would use it up exactly at the end of the...

Writing Recording Rules for SLIs

Why raw long-window queries hurt Running a query like rate(...[3d]) on every alert evaluation makes Prometheus re-read three days of data every...

Skills You'll Master

SLOERROR-BUDGETBURN-RATE-ALERTSPROMETHEUSSRE

Curriculum Index12 topics

1

Understanding Why Dashboards Are Not Enough

It is the Friday of acme-shop's biggest sale.

2

Choosing SLIs That Reflect User Pain

What an SLI is and the three common shapes An SLI (Service Level Indicator) is one precisely defined measurement of...

3

Setting SLOs That Match What Users Need

What an SLO is An SLO (Service Level Objective) is the target value for an SLI over a defined time window.

4

Spending and Protecting the Error Budget

What an error budget is The error budget is the unreliability your SLO allows: one minus the SLO, applied to the events...

5

Understanding Burn Rate and Why One Threshold Fails

What burn rate means Burn rate is how fast you are spending the error budget compared with the pace that would use it...

6

Writing Recording Rules for SLIs

Why raw long-window queries hurt Running a query like rate(...[3d]) on every alert evaluation makes Prometheus re-read...

7

Writing Multi-Window Burn-Rate Alerts

The three alert tiers A good starting setup uses three tiers: two that page a human and one that opens a ticket.

8

Reading the Four Golden Signals

The four signals The four golden signals are latency, traffic, errors, and saturation, and together they describe the...

9

Connecting Logs and Traces to an SLO Breach

Adding correlation IDs to logs A correlation ID is a unique value attached to a request at the edge and written to...

10

Designing SLO Dashboards and Routing Pages

Building different dashboards for different readers Different readers need different dashboards.

11

Building SLOs and Alerts in a Hands-on Lab

Preparing the sre-lab cluster You will give checkout-service two SLOs, then break it in two different ways and watch...

12

Quick Reference and Common Mistakes

Quick reference Keep these formulas and defaults close when you design an SLO.

Frequently Asked Questions

An SLI is a measurement, such as the share of requests that succeed. An SLO is the internal target for that measurement, such as 99.9 percent over 30 days. An SLA is a contract with customers, usually looser than the SLO, with penalties if it is broken.

It is the amount of unreliability your SLO allows. A 99.9 percent SLO leaves 0.1 percent of requests free to fail. Teams spend that budget on releases and experiments, and slow down when it runs low.

A long window proves enough budget has really been spent. A short window proves the problem is still happening now. Requiring both stops pages for blips that already recovered, and still catches slow leaks.

Start with one availability SLO and one latency SLO on the user-facing request path, set a little looser than current performance. You can tighten the targets later, once you trust the measurements.