Skip to main content

Chaos Engineering for SRE

Learn to break systems on purpose, safely: steady-state hypotheses, blast radius, Chaos Mesh experiments, Game Days, and FMEA to pick what to test.

~3 hours
8 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Chaos Engineering Exists

acme-shop's checkout-service runs three replicas. Dashboards are green, and everyone assumes that if one pod dies, the other two absorb the traffic.

Writing a Steady-State Hypothesis

You cannot judge an experiment without a precise definition of "working". That definition is the steady-state hypothesis.

Controlling Blast Radius and Stopping Safely

A chaos experiment is only as safe as its limits. Decide them before you start, not while the dashboard is turning red.

Running Game Days

Experiments test the system. A Game Day tests the system and the humans together.

Choosing What to Test With FMEA

You cannot test everything at once. A simple FMEA worksheet tells you where to begin.

Running Experiments With Chaos Mesh

Chaos Mesh lets you describe a fault as a Kubernetes resource, so experiments are reviewable, repeatable, and time-limited.

Skills You'll Master

CHAOS-ENGINEERINGCHAOS-MESHGAME-DAYFMEASRE

Curriculum Index8 topics

Career Impact

Roles that use the skills in this module.

  • Site Reliability Engineer

  • DevOps Engineer

  • Platform Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

It can be, once you have a written hypothesis, a small blast radius, and a way to stop immediately. Teams build up to it by starting in staging. Running a first experiment directly in production without that track record is how experiments turn into incidents.

It is a measurable statement of what normal looks like, written before the experiment. For example, p99 latency stays below 500 ms and errors stay below 0.5 percent while one pod is killed. If you cannot measure it, you cannot tell whether the experiment passed.

An experiment tests the system. A Game Day tests the system and the people together, running the real incident process against a planned failure. Both are useful, and a Game Day often reveals broken runbooks that an experiment never would.

No. Killing a pod by hand is a valid first experiment. Chaos Mesh helps when you want repeatable, declarative, time-limited faults such as network delay, which are hard to inject safely by hand.