Databases and Storage Reliability for SRE
Learn to keep PostgreSQL and Redis reliable: WAL, vacuum, connection pooling, replication lag, failover, point-in-time recovery, and stateful storage.
What You'll Learn
Understanding Why Databases Break SREs Differently
By now acme-shop's checkout-service survives a dead recommendations service and ships behind feature flags.
Understanding the PostgreSQL Write-Ahead Log
The write-ahead log is the one mechanism behind crash recovery, replication, and point-in-time recovery.
Keeping Vacuum Healthy and Avoiding Wraparound
Vacuum is the quiet background job whose absence causes a loud outage months later.
Pooling Connections With PgBouncer
A connection pool is the cheapest way to stop 200 pods from each opening their own handful of database connections.
Measuring Replication Lag and Failover
Replicas give you read capacity and a failover target, but only if you know how far behind they are.
Recovering to a Point in Time
Point-in-time recovery (PITR) is how you undo a bad migration without losing the rest of the day.
Skills You'll Master
Curriculum Index11 topics
Understanding Why Databases Break SREs Differently
By now acme-shop's checkout-service survives a dead recommendations service and ships behind feature flags.
Understanding the PostgreSQL Write-Ahead Log
The write-ahead log is the one mechanism behind crash recovery, replication, and point-in-time recovery.
Keeping Vacuum Healthy and Avoiding Wraparound
Vacuum is the quiet background job whose absence causes a loud outage months later.
Pooling Connections With PgBouncer
A connection pool is the cheapest way to stop 200 pods from each opening their own handful of database connections.
Measuring Replication Lag and Failover
Replicas give you read capacity and a failover target, but only if you know how far behind they are.
Recovering to a Point in Time
Point-in-time recovery (PITR) is how you undo a bad migration without losing the rest of the day.
Running PostgreSQL on Kubernetes With CloudNativePG
Operators exist because promoting a replica, re-pointing traffic, and rebuilding a failed instance is repetitive work...
Keeping Redis Safe: Persistence, Eviction, and Availability
Redis is treated as a cache until the day it holds something that cannot be rebuilt.
Choosing Database SLIs and Alerts
You already know how to write SLOs. The skill here is choosing database signals that predict user pain.
Building the Hands-On Database Reliability Lab
You will run PostgreSQL with a replica, exhaust a pool, watch a stale read, kill the primary, and test Redis memory...
Quick Reference and Common Mistakes
Quick reference Common mistakes Treating an idle database as an innocent database is the most expensive assumption in...
Career Impact
Roles that use the skills in this module.
Site Reliability Engineer
DevOps Engineer
Platform Engineer
Next Modules
Related Guides
Practice on the Coding Sheet
Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.
Open the Coding SheetFrequently Asked Questions
The bottleneck is often in front of the database, not inside it. A full connection pool, a lagging replica, or a blocked vacuum can all make requests time out while database CPU stays idle. Check pool utilisation and replication lag before you tune queries.
A base backup is a copy of the data files at one moment. The write-ahead log (WAL) is the stream of every change after that moment. You need both to restore to an exact time.
For most stateless web services, yes. Transaction mode returns a server connection to the pool after each transaction, so a few connections serve many clients. Use session mode only if your code depends on session features such as advisory locks.
It can be either, and the difference decides your settings. If every key can be rebuilt from PostgreSQL, treat it as a cache and let it evict. If it holds data you cannot rebuild, turn on AOF persistence and plan for failover.