Distributed Systems Failure Modes for SRE
Learn why distributed systems fail in production: CAP, consensus, cascading failures, retry storms, and the patterns that contain them, with labs.
What You'll Learn
Understanding Why Distributed Systems Fail Differently
You have run your first acme-shop incident, and it had a clear cause: a bad release. The next one is stranger.
Applying the CAP Theorem to Real Database Choices
When engineers hear "CAP theorem" they think it is academic. It is not. It is one of the most practical lenses for choosing a database or a queue.
Understanding How Raft Achieves Consensus
Many systems need several machines to agree on one thing, even when some fail.
Diagnosing Cascading Failures Before They Spread
A cascading failure is the main way small problems become large outages. Learn its shape and you will recognise it early.
Building Resilience Patterns That Actually Hold Under Load
Now the fixes. Each pattern below limits how far one slow or failing component can spread.
Understanding Head-of-Line Blocking
What it is Head-of-line blocking happens when one slow request holds up other, unrelated requests queued behind it on the same connection or queue.
Skills You'll Master
Curriculum Index9 topics
Understanding Why Distributed Systems Fail Differently
You have run your first acme-shop incident, and it had a clear cause: a bad release. The next one is stranger.
Applying the CAP Theorem to Real Database Choices
When engineers hear "CAP theorem" they think it is academic. It is not.
Understanding How Raft Achieves Consensus
Many systems need several machines to agree on one thing, even when some fail.
Diagnosing Cascading Failures Before They Spread
A cascading failure is the main way small problems become large outages.
Building Resilience Patterns That Actually Hold Under Load
Now the fixes. Each pattern below limits how far one slow or failing component can spread.
Understanding Head-of-Line Blocking
What it is Head-of-line blocking happens when one slow request holds up other, unrelated requests queued behind it on...
Hands-On Lab: Building and Breaking Resilience Patterns
📌 Remember: This lab costs nothing. It needs Python 3.10 or later on sre-vm (open it with multipass shell sre-vm) or...
Quick Reference
Common mistakes engineers make with distributed systems Treating a timeout as a failure and retrying at once, with no...
What You Built and What Comes Next
You traced how one slow payment-service call can take down a whole shop, and you saw why slow is worse than down.
Career Impact
Roles that use the skills in this module.
Site Reliability Engineer
DevOps Engineer
Platform Engineer
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
A down service fails fast, so callers get an error and move on. A slow service keeps threads and connections busy for far longer than planned. Those resources run out across the whole call graph, and the slowness spreads.
You do not need to implement it, but you do need to know why etcd, Consul, and ZooKeeper stop accepting writes when they lose a majority. That behaviour is by design and protects you from split-brain.
Jitter adds a random amount to each retry delay. Without it, thousands of clients that failed together retry together and knock the recovering service down again. Random delays spread that load over time.
No. It is a small teaching version that shows the state machine. In production, use a mature library or the circuit breaking built into your service mesh, which handles failure rates, concurrency, and metrics.