Skip to main content

Incident Management and On-Call for SRE

Learn to run incidents calmly: severity levels, the incident commander role, a six-step debugging method, blameless postmortems, and healthy on-call.

~3 hours
11 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Incidents Go Wrong Before Anyone Touches Code

You joined acme-shop as its first SRE. Checkout now has SLOs and burn-rate alerts, and tonight one of those alerts fires for real.

Understanding Incident Severity Levels

Not every alert is a 2 AM phone call. A severity system lets engineering, support, and leadership agree on how urgently to respond, without arguing...

Understanding the Incident Commander Role

The incident commander (IC) is the one person who owns coordination during an incident.

Understanding the First Five Minutes

The first five minutes decide whether the incident stays small. The instinct is to start fixing immediately. Resist it.

Understanding Communication During an Incident

People who are not debugging still need to know what is going on. Without a plan, they interrupt the people who are.

Understanding the Six-Step Debugging Method

This is the core operational skill of the module. Every real incident, whatever the technology, follows the same six steps.

Skills You'll Master

INCIDENT-MANAGEMENTON-CALLPOSTMORTEMINCIDENT-RESPONSESRE

Curriculum Index11 topics

1

Understanding Why Incidents Go Wrong Before Anyone Touches Code

You joined acme-shop as its first SRE. Checkout now has SLOs and burn-rate alerts, and tonight one of those alerts...

2

Understanding Incident Severity Levels

Not every alert is a 2 AM phone call. A severity system lets engineering, support, and leadership agree on how urgently...

3

Understanding the Incident Commander Role

The incident commander (IC) is the one person who owns coordination during an incident.

4

Understanding the First Five Minutes

The first five minutes decide whether the incident stays small. The instinct is to start fixing immediately. Resist it.

5

Understanding Communication During an Incident

People who are not debugging still need to know what is going on. Without a plan, they interrupt the people who are.

6

Understanding the Six-Step Debugging Method

This is the core operational skill of the module.

7

Understanding Blameless Postmortems

The incident is over, but the work is not. A postmortem turns one bad night into a lasting improvement.

8

Understanding On-Call Health

A team can run perfect incident response and still burn out if on-call itself is unhealthy.

9

Building the Hands-On Incident Response Lab

📌 Remember: This lab runs entirely on your laptop in the sre-lab kind cluster, so it costs nothing.

10

Quick Reference and Common Mistakes

Quick reference Common mistakes Changing things before you know the blast radius is the fastest way to turn one...

11

What You Built and What Comes Next

You ran your first acme-shop incident from alert to verified fix.

Career Impact

Roles that use the skills in this module.

  • Site Reliability Engineer

  • DevOps Engineer

  • Platform Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

The incident commander coordinates the response: who is working on what, what has been tried, and what happens next. They do not have to fix the bug. Splitting coordination from fixing keeps the people debugging focused.

A restart erases evidence such as memory growth, thread dumps, and error patterns. It can also hide the symptom for a few minutes while the real cause keeps working. Read the metrics first, then decide what to change.

A blameless postmortem looks for the system, process, and tooling gaps that let a mistake cause impact, instead of naming a person as the cause. People then describe what really happened, including wrong turns, which is the information the team needs.

Track pages per shift, the share of pages that needed no action, and pages outside working hours. If engineers ignore pages or dread shifts, the rotation is unhealthy even when incident response looks good on paper.