Incident Management and On-Call for SRE
Learn to run incidents calmly: severity levels, the incident commander role, a six-step debugging method, blameless postmortems, and healthy on-call.
What You'll Learn
Understanding Why Incidents Go Wrong Before Anyone Touches Code
You joined acme-shop as its first SRE. Checkout now has SLOs and burn-rate alerts, and tonight one of those alerts fires for real.
Understanding Incident Severity Levels
Not every alert is a 2 AM phone call. A severity system lets engineering, support, and leadership agree on how urgently to respond, without arguing...
Understanding the Incident Commander Role
The incident commander (IC) is the one person who owns coordination during an incident.
Understanding the First Five Minutes
The first five minutes decide whether the incident stays small. The instinct is to start fixing immediately. Resist it.
Understanding Communication During an Incident
People who are not debugging still need to know what is going on. Without a plan, they interrupt the people who are.
Understanding the Six-Step Debugging Method
This is the core operational skill of the module. Every real incident, whatever the technology, follows the same six steps.
Skills You'll Master
Curriculum Index11 topics
Understanding Why Incidents Go Wrong Before Anyone Touches Code
You joined acme-shop as its first SRE. Checkout now has SLOs and burn-rate alerts, and tonight one of those alerts...
Understanding Incident Severity Levels
Not every alert is a 2 AM phone call. A severity system lets engineering, support, and leadership agree on how urgently...
Understanding the Incident Commander Role
The incident commander (IC) is the one person who owns coordination during an incident.
Understanding the First Five Minutes
The first five minutes decide whether the incident stays small. The instinct is to start fixing immediately. Resist it.
Understanding Communication During an Incident
People who are not debugging still need to know what is going on. Without a plan, they interrupt the people who are.
Understanding the Six-Step Debugging Method
This is the core operational skill of the module.
Understanding Blameless Postmortems
The incident is over, but the work is not. A postmortem turns one bad night into a lasting improvement.
Understanding On-Call Health
A team can run perfect incident response and still burn out if on-call itself is unhealthy.
Building the Hands-On Incident Response Lab
📌 Remember: This lab runs entirely on your laptop in the sre-lab kind cluster, so it costs nothing.
Quick Reference and Common Mistakes
Quick reference Common mistakes Changing things before you know the blast radius is the fastest way to turn one...
What You Built and What Comes Next
You ran your first acme-shop incident from alert to verified fix.
Career Impact
Roles that use the skills in this module.
Site Reliability Engineer
DevOps Engineer
Platform Engineer
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
The incident commander coordinates the response: who is working on what, what has been tried, and what happens next. They do not have to fix the bug. Splitting coordination from fixing keeps the people debugging focused.
A restart erases evidence such as memory growth, thread dumps, and error patterns. It can also hide the symptom for a few minutes while the real cause keeps working. Read the metrics first, then decide what to change.
A blameless postmortem looks for the system, process, and tooling gaps that let a mistake cause impact, instead of naming a person as the cause. People then describe what really happened, including wrong turns, which is the information the team needs.
Track pages per shift, the share of pages that needed no action, and pages outside working hours. If engineers ignore pages or dread shifts, the rotation is unhealthy even when incident response looks good on paper.