Skip to main content

Databases and Storage Reliability for SRE

Learn to keep PostgreSQL and Redis reliable: WAL, vacuum, connection pooling, replication lag, failover, point-in-time recovery, and stateful storage.

~3.5 hours
11 Topics
Hands-on Scenarios

What You'll Learn

Understanding Why Databases Break SREs Differently

By now acme-shop's checkout-service survives a dead recommendations service and ships behind feature flags.

Understanding the PostgreSQL Write-Ahead Log

The write-ahead log is the one mechanism behind crash recovery, replication, and point-in-time recovery.

Keeping Vacuum Healthy and Avoiding Wraparound

Vacuum is the quiet background job whose absence causes a loud outage months later.

Pooling Connections With PgBouncer

A connection pool is the cheapest way to stop 200 pods from each opening their own handful of database connections.

Measuring Replication Lag and Failover

Replicas give you read capacity and a failover target, but only if you know how far behind they are.

Recovering to a Point in Time

Point-in-time recovery (PITR) is how you undo a bad migration without losing the rest of the day.

Skills You'll Master

POSTGRESQLREDISDATABASE-RELIABILITYCLOUDNATIVEPGSRE

Curriculum Index11 topics

1

Understanding Why Databases Break SREs Differently

By now acme-shop's checkout-service survives a dead recommendations service and ships behind feature flags.

2

Understanding the PostgreSQL Write-Ahead Log

The write-ahead log is the one mechanism behind crash recovery, replication, and point-in-time recovery.

3

Keeping Vacuum Healthy and Avoiding Wraparound

Vacuum is the quiet background job whose absence causes a loud outage months later.

4

Pooling Connections With PgBouncer

A connection pool is the cheapest way to stop 200 pods from each opening their own handful of database connections.

5

Measuring Replication Lag and Failover

Replicas give you read capacity and a failover target, but only if you know how far behind they are.

6

Recovering to a Point in Time

Point-in-time recovery (PITR) is how you undo a bad migration without losing the rest of the day.

7

Running PostgreSQL on Kubernetes With CloudNativePG

Operators exist because promoting a replica, re-pointing traffic, and rebuilding a failed instance is repetitive work...

8

Keeping Redis Safe: Persistence, Eviction, and Availability

Redis is treated as a cache until the day it holds something that cannot be rebuilt.

9

Choosing Database SLIs and Alerts

You already know how to write SLOs. The skill here is choosing database signals that predict user pain.

10

Building the Hands-On Database Reliability Lab

You will run PostgreSQL with a replica, exhaust a pool, watch a stale read, kill the primary, and test Redis memory...

11

Quick Reference and Common Mistakes

Quick reference Common mistakes Treating an idle database as an innocent database is the most expensive assumption in...

Career Impact

Roles that use the skills in this module.

  • Site Reliability Engineer

  • DevOps Engineer

  • Platform Engineer

See how this is asked in interviews

Practice on the Coding Sheet

Not a software engineer sheet. Every problem comes from real DevOps, SRE, Platform and Cloud interviews, from your first script to a system you build yourself.

Open the Coding Sheet

Frequently Asked Questions

The bottleneck is often in front of the database, not inside it. A full connection pool, a lagging replica, or a blocked vacuum can all make requests time out while database CPU stays idle. Check pool utilisation and replication lag before you tune queries.

A base backup is a copy of the data files at one moment. The write-ahead log (WAL) is the stream of every change after that moment. You need both to restore to an exact time.

For most stateless web services, yes. Transaction mode returns a server connection to the pool after each transaction, so a few connections serve many clients. Use session mode only if your code depends on session features such as advisory locks.

It can be either, and the difference decides your settings. If every key can be rebuilt from PostgreSQL, treat it as a cache and let it evict. If it holds data you cannot rebuild, turn on AOF persistence and plan for failover.