SLOs and Observability for SRE
Learn to define SLIs and SLOs, track error budgets, and write multi-window burn-rate alerts that page only when users are really affected.
What You'll Learn
Understanding Why Dashboards Are Not Enough
It is the Friday of acme-shop's biggest sale. Every Grafana panel is green: CPU at 40 percent, every pod Running, no restarts.
Choosing SLIs That Reflect User Pain
What an SLI is and the three common shapes An SLI (Service Level Indicator) is one precisely defined measurement of something users experience...
Setting SLOs That Match What Users Need
What an SLO is An SLO (Service Level Objective) is the target value for an SLI over a defined time window.
Spending and Protecting the Error Budget
What an error budget is The error budget is the unreliability your SLO allows: one minus the SLO, applied to the events in the window.
Understanding Burn Rate and Why One Threshold Fails
What burn rate means Burn rate is how fast you are spending the error budget compared with the pace that would use it up exactly at the end of the...
Writing Recording Rules for SLIs
Why raw long-window queries hurt Running a query like rate(...[3d]) on every alert evaluation makes Prometheus re-read three days of data every...
Skills You'll Master
Curriculum Index12 topics
Understanding Why Dashboards Are Not Enough
It is the Friday of acme-shop's biggest sale.
Choosing SLIs That Reflect User Pain
What an SLI is and the three common shapes An SLI (Service Level Indicator) is one precisely defined measurement of...
Setting SLOs That Match What Users Need
What an SLO is An SLO (Service Level Objective) is the target value for an SLI over a defined time window.
Spending and Protecting the Error Budget
What an error budget is The error budget is the unreliability your SLO allows: one minus the SLO, applied to the events...
Understanding Burn Rate and Why One Threshold Fails
What burn rate means Burn rate is how fast you are spending the error budget compared with the pace that would use it...
Writing Recording Rules for SLIs
Why raw long-window queries hurt Running a query like rate(...[3d]) on every alert evaluation makes Prometheus re-read...
Writing Multi-Window Burn-Rate Alerts
The three alert tiers A good starting setup uses three tiers: two that page a human and one that opens a ticket.
Reading the Four Golden Signals
The four signals The four golden signals are latency, traffic, errors, and saturation, and together they describe the...
Connecting Logs and Traces to an SLO Breach
Adding correlation IDs to logs A correlation ID is a unique value attached to a request at the edge and written to...
Designing SLO Dashboards and Routing Pages
Building different dashboards for different readers Different readers need different dashboards.
Building SLOs and Alerts in a Hands-on Lab
Preparing the sre-lab cluster You will give checkout-service two SLOs, then break it in two different ways and watch...
Quick Reference and Common Mistakes
Quick reference Keep these formulas and defaults close when you design an SLO.
Career Impact
Roles that use the skills in this module.
Site Reliability Engineer
DevOps Engineer
Platform Engineer
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
An SLI is a measurement, such as the share of requests that succeed. An SLO is the internal target for that measurement, such as 99.9 percent over 30 days. An SLA is a contract with customers, usually looser than the SLO, with penalties if it is broken.
It is the amount of unreliability your SLO allows. A 99.9 percent SLO leaves 0.1 percent of requests free to fail. Teams spend that budget on releases and experiments, and slow down when it runs low.
A long window proves enough budget has really been spent. A short window proves the problem is still happening now. Requiring both stops pages for blips that already recovered, and still catches slow leaks.
Start with one availability SLO and one latency SLO on the user-facing request path, set a little looser than current performance. You can tighten the targets later, once you trust the measurements.