Skip to main content

Running StatefulSets for Databases on Kubernetes

Deploy and manage stateful database workloads on Kubernetes using StatefulSets with stable network identities, ordered scaling, and persistent storage.

52 Terms

Overview and What You Will Learn

Regular Deployments treat every pod as identical and interchangeable — perfect for stateless APIs but catastrophic for databases where pod identity, startup order, and storage persistence are critical. StatefulSets solve this by giving each pod a stable, predictable identity, its own dedicated PersistentVolumeClaim, and strict ordered deployment and termination guarantees. This lab walks you through deploying PostgreSQL, Redis, and a multi-node database cluster on Kubernetes using StatefulSets with production-grade configuration.

By the end of this guide you will be able to:

  • Understand the core differences between Deployments and StatefulSets and when to use each
  • Deploy a single-instance PostgreSQL database using a StatefulSet with persistent storage
  • Configure a Redis cluster using StatefulSets with stable network identities
  • Set up a primary-replica PostgreSQL configuration with ordered pod startup
  • Troubleshoot common StatefulSet failures including stuck termination and PVC binding issues

Why This Matters in Production

Zerodha runs PostgreSQL for trade records and MySQL for user accounts directly on Kubernetes using StatefulSets. The ordered startup guarantee means the primary database pod always initialises and becomes ready before replica pods attempt to connect and begin replication — preventing the split-brain scenarios that plague manually managed database clusters.

At Razorpay, Redis is deployed as a StatefulSet cluster where each node has a stable DNS name (redis-0.redis, redis-1.redis, redis-2.redis) that never changes even after pod restarts. Application code hardcodes these stable names rather than dynamic pod IPs — impossible with a regular Deployment.

Core Principles

StatefulSet vs Deployment — the critical differences: DEPLOYMENT STATEFULSET ────────── ─────────── Pod names Random suffix Stable ordinal api-7d9f8b-xkp2q postgres-0 api-7d9f8b-mn3lp postgres-1 postgres-2 Pod identity Interchangeable Unique and stable Storage Shared or none Each pod gets its own dedicated PVC (postgres-data-0, postgres-data-1) Startup order All pods start Ordered: pod-0 must simultaneously be Ready before pod-1 starts Termination All pods stop Reverse order: simultaneously pod-2 → pod-1 → pod-0 DNS Service IP only Per-pod DNS: pod-0.service.ns.svc.cluster.local

When to use StatefulSet vs Deployment: Use StatefulSet when: Use Deployment when: ──────────────────── ──────────────────

Databases (PostgreSQL, MySQL) * REST APIs Message queues (Kafka, RabbitMQ) * Web servers (NGINX, Express) Caches with persistence (Redis) * Background workers (stateless) Search engines (Elasticsearch) * Any app with no local state Any app needing stable pod DNS * Any app that is truly stateless

Detailed Step-by-Step Practical Lab

Step 1 — Create the Headless Service for Stable Pod DNS

StatefulSets require a Headless Service — a Service with clusterIP: None that creates individual DNS entries for each pod instead of a single load-balanced IP:

YAML
apiVersion: v1
kind: Service
metadata:
name: postgres
namespace: production
labels:
app: postgres
spec:
clusterIP: None # This makes it a Headless Service
selector:
app: postgres
ports:
- name: postgres
port: 5432
targetPort: 5432
Bash
kubectl apply -f headless-service-postgres.yaml
# This creates DNS entries for each pod:
# postgres-0.postgres.production.svc.cluster.local → pod IP of postgres-0
# postgres-1.postgres.production.svc.cluster.local → pod IP of postgres-1
# postgres-2.postgres.production.svc.cluster.local → pod IP of postgres-2
# Also create a regular Service for client connections (load balances reads)
kubectl apply -f - <<EOF
apiVersion: v1
kind: Service
metadata:
name: postgres-primary
namespace: production
spec:
selector:
app: postgres
role: primary # Only route to the primary pod
ports:
- port: 5432
targetPort: 5432
EOF
Remember

The Headless Service name must match the serviceName field in your StatefulSet spec — this is what enables the stable per-pod DNS names. Getting this wrong is the most common StatefulSet configuration mistake.

Step 2 — Deploy Single-Instance PostgreSQL StatefulSet
YAML
# statefulset-postgres.yaml — production PostgreSQL on Kubernetes
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: postgres
namespace: production
spec:
serviceName: "postgres" # Must match the Headless Service name
replicas: 1 # Start single — add replicas for HA
selector:
matchLabels:
app: postgres
template:
metadata:
labels:
app: postgres
role: primary
spec:
terminationGracePeriodSeconds: 60 # Give PostgreSQL time to flush WAL
securityContext:
fsGroup: 999 # postgres UID — sets volume ownership
runAsUser: 999
runAsNonRoot: true
initContainers:
# Fix permissions on the data directory before PostgreSQL starts
- name: fix-permissions
image: busybox:1.35
command: ["sh", "-c", "chown -R 999:999 /var/lib/postgresql/data"]
volumeMounts:
- name: postgres-data
mountPath: /var/lib/postgresql/data
securityContext:
runAsUser: 0 # Run as root for chown only
containers:
- name: postgres
image: postgres:15.4
ports:
- containerPort: 5432
name: postgres
env:
- name: POSTGRES_DB
value: "zerodha_trading"
- name: POSTGRES_USER
valueFrom:
secretKeyRef:
name: postgres-credentials
key: username
- name: POSTGRES_PASSWORD
valueFrom:
secretKeyRef:
name: postgres-credentials
key: password
- name: PGDATA
value: "/var/lib/postgresql/data/pgdata" # Subdirectory avoids lost+found
- name: POSTGRES_INITDB_ARGS
value: "--encoding=UTF8 --auth-host=scram-sha-256"
resources:
requests:
cpu: "500m"
memory: "1Gi"
limits:
cpu: "4"
memory: "8Gi"
livenessProbe:
exec:
command:
- pg_isready
- -U
- $(POSTGRES_USER)
- -d
- $(POSTGRES_DB)
initialDelaySeconds: 30
periodSeconds: 10
failureThreshold: 6
readinessProbe:
exec:
command:
- pg_isready
- -U
- $(POSTGRES_USER)
- -d
- $(POSTGRES_DB)
initialDelaySeconds: 5
periodSeconds: 5
failureThreshold: 3
volumeMounts:
- name: postgres-data
mountPath: /var/lib/postgresql/data
- name: postgres-config
mountPath: /etc/postgresql/postgresql.conf
subPath: postgresql.conf
volumeClaimTemplates: # Each pod gets its own PVC automatically
- metadata:
name: postgres-data
labels:
app: postgres
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: gp3-encrypted
resources:
requests:
storage: 100Gi
Bash
kubectl apply -f statefulset-postgres.yaml
# Watch ordered pod startup
kubectl get pods -n production -w
# NAME READY STATUS RESTARTS
# postgres-0 0/1 ContainerCreating 0 ← starts first
# postgres-0 0/1 Running 0
# postgres-0 1/1 Running 0 ← must be Ready before replicas start
# Verify PVC was automatically created
kubectl get pvc -n production
# NAME STATUS VOLUME CAPACITY
# postgres-data-postgres-0 Bound pvc-a1b2c3d4-... 100Gi
Step 3 — Deploy Redis as a StatefulSet Cluster
YAML
# statefulset-redis.yaml — Redis cluster with stable pod identities
apiVersion: v1
kind: ConfigMap
metadata:
name: redis-config
namespace: production
data:
redis.conf: |
maxmemory 2gb
maxmemory-policy allkeys-lru
appendonly yes
appendfsync everysec
save 900 1
save 300 10
save 60 10000
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: redis
namespace: production
spec:
serviceName: "redis"
replicas: 3 # 3-node Redis cluster
selector:
matchLabels:
app: redis
template:
metadata:
labels:
app: redis
spec:
terminationGracePeriodSeconds: 30
containers:
- name: redis
image: redis:7.2
command: ["redis-server", "/etc/redis/redis.conf"]
ports:
- containerPort: 6379
name: redis
resources:
requests:
cpu: "250m"
memory: "512Mi"
limits:
cpu: "1"
memory: "2Gi"
livenessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 15
periodSeconds: 10
readinessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 5
periodSeconds: 5
volumeMounts:
- name: redis-data
mountPath: /data
- name: redis-config
mountPath: /etc/redis
volumes:
- name: redis-config
configMap:
name: redis-config
volumeClaimTemplates:
- metadata:
name: redis-data
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: gp3-encrypted
resources:
requests:
storage: 20Gi
Bash
kubectl apply -f statefulset-redis.yaml
# Watch all 3 Redis pods start in strict order
kubectl get pods -n production -w
# redis-0 1/1 Running 0 ← starts and becomes Ready first
# redis-1 1/1 Running 0 ← starts only after redis-0 is Ready
# redis-2 1/1 Running 0 ← starts only after redis-1 is Ready
# Connect to Redis and verify cluster
kubectl exec -it redis-0 -n production -- redis-cli ping
# PONG
# Each pod has a stable DNS name — application connects using these
# redis-0.redis.production.svc.cluster.local:6379
# redis-1.redis.production.svc.cluster.local:6379
# redis-2.redis.production.svc.cluster.local:6379
Step 4 — Perform a Rolling Update on a StatefulSet
Bash
# Update PostgreSQL image version
kubectl set image statefulset/postgres \
postgres=postgres:15.5 \
-n production
# Watch ordered rolling update — updates in reverse order (pod-2 first, pod-0 last)
kubectl rollout status statefulset/postgres -n production
# Waiting for 1 pods to be ready...
# statefulset rolling update complete 1 pods at revision postgres-6d8f9b...
# Check rollout history
kubectl rollout history statefulset/postgres -n production
# Rollback if needed
kubectl rollout undo statefulset/postgres -n production
Tip

StatefulSet rolling updates go in reverse ordinal order — pod-2 is updated first, then pod-1, then pod-0. For primary-replica databases this means replicas are updated before the primary, which is the safe order. Always verify replication lag is zero before each pod update completes.

Step 5 — Scale a StatefulSet Up and Down Safely
Bash
# Scale up — new pods start in order (pod-1 after pod-0 is Ready)
kubectl scale statefulset postgres -n production --replicas=3
# Watch ordered scale-up
kubectl get pods -n production -w
# postgres-0 1/1 Running 0
# postgres-1 0/1 Pending 0 ← starts after postgres-0 is Ready
# postgres-1 1/1 Running 0
# postgres-2 0/1 Pending 0 ← starts after postgres-1 is Ready
# postgres-2 1/1 Running 0
# Scale down — pods terminate in reverse order (pod-2 first)
kubectl scale statefulset postgres -n production --replicas=1
# CRITICAL: Scaling down does NOT delete PVCs
# PVCs for postgres-1 and postgres-2 still exist after scale-down
kubectl get pvc -n production | grep postgres
# postgres-data-postgres-0 Bound 100Gi ← active
# postgres-data-postgres-1 Bound 100Gi ← orphaned — delete manually if not needed
# postgres-data-postgres-2 Bound 100Gi ← orphaned — delete manually if not needed
Security

Never delete orphaned PVCs automatically. Kubernetes intentionally keeps them to prevent accidental data loss. Review and manually delete them only after confirming the data is either replicated elsewhere or no longer needed.

Step 6 — Troubleshoot Common StatefulSet Failures
Bash
# Problem 1 — Pod stuck in Terminating state
kubectl get pods -n production
# postgres-0 1/1 Terminating 0 48m ← stuck
# Cause: The pod has a finalizer or the node is unresponsive
# Check for finalizers
kubectl get pod postgres-0 -n production -o jsonpath='{.metadata.finalizers}'
# Force delete as last resort (data loss risk — only if node is dead)
kubectl delete pod postgres-0 -n production --force --grace-period=0
# Problem 2 — PVC stuck in Pending after scale-up
kubectl describe pvc postgres-data-postgres-1 -n production
# Events: ProvisioningFailed: no nodes available in zone ap-south-1a
# Cause: WaitForFirstConsumer mode — pod must be scheduled first
# Fix: Ensure the pod is scheduled before checking PVC status
# Problem 3 — Pod-1 stuck in Init state waiting for pod-0
kubectl get pods -n production
# postgres-0 0/1 Running 0 ← not Ready yet (probe failing)
# postgres-1 0/1 Init:0/1 0 ← waiting for postgres-0 to be Ready
# Check why postgres-0 is not passing readiness probe
kubectl describe pod postgres-0 -n production
kubectl logs postgres-0 -n production

Production Best Practices & Common Pitfalls

  • Always set terminationGracePeriodSeconds to at least 60 for databases. The default 30 seconds is too short for PostgreSQL to complete a checkpoint and flush WAL — abrupt termination risks data corruption.
  • Use podManagementPolicy: Parallel only for StatefulSets where pods are truly independent — like Elasticsearch data nodes. Never use it for primary-replica databases where order matters.
  • Monitor replication lag on all replica pods. A replica that falls too far behind the primary will cause data loss if the primary fails before the replica catches up.
  • Back up PVCs using Velero with volume snapshots on a schedule — at minimum daily, ideally every hour for financial transaction databases.
  • Use updateStrategy: RollingUpdate with partition during major database version upgrades — this lets you upgrade one pod at a time and pause the rollout to verify replication before continuing.
Common Mistake

Deleting a StatefulSet with kubectl delete statefulset postgres thinking it will also clean up PVCs. It does not — PVCs are intentionally orphaned. But the pods are deleted, leaving your database inaccessible until the StatefulSet is recreated and the pods rebind to the orphaned PVCs. Always scale to zero first, verify, then delete.

Quick Reference & Troubleshooting Commands

Command Purpose
kubectl get statefulset -n <ns> List all StatefulSets and replica counts
kubectl describe statefulset <name> -n <ns> Full StatefulSet config and events
kubectl get pods -n <ns> -w Watch ordered pod startup and termination
kubectl scale statefulset <name> --replicas=<n> -n <ns> Scale StatefulSet up or down
kubectl rollout status statefulset <name> -n <ns> Watch rolling update progress
kubectl rollout undo statefulset <name> -n <ns> Rollback to previous StatefulSet revision
kubectl exec -it <name>-0 -n <ns> -- bash Shell into the primary pod (ordinal 0)
kubectl get pvc -n <ns> | grep <statefulset-name> List PVCs created by a StatefulSet
kubectl delete pod <name>-0 -n <ns> --force --grace-period=0 Force delete stuck Terminating pod
kubectl get pod <name>-0 -n <ns> -o jsonpath='{.metadata.finalizers}' Check for blocking finalizers

Resources

Velero vs etcd Snapshot: Not a Real Choice

Velero vs etcd Snapshot: Not a Real Choice

Velero and etcd snapshots protect different layers of a cluster, not the same thing. Here's what each covers and why production DR needs both.

5 min read•Aug 2026
Istio Ambient vs Linkerd in 2026

Istio Ambient vs Linkerd in 2026

Istio Ambient killed the sidecar-tax argument. The real 2026 decision is waypoint topology and Buoyant's licensing shift, not features vs simplicity.

5 min read•Aug 2026
Nginx Ingress vs Traefik vs Gateway API in 2026

Nginx Ingress vs Traefik vs Gateway API in 2026

Ingress-nginx retired in March 2026. Here's how Traefik and the Gateway API actually compare as replacements — and why "just swap it" is the wrong frame.

5 min read•Aug 2026
OPA Gatekeeper vs Kyverno: Policy Engine in 2026

OPA Gatekeeper vs Kyverno: Policy Engine in 2026

OPA Gatekeeper vs Kyverno compared for 2026 - Rego vs YAML, mutation maturity, operational overhead, and which policy engine fits your cluster.

5 min read•Aug 2026
Helm vs Kustomize in 2026: Templating vs Patching

Helm vs Kustomize in 2026: Templating vs Patching

Helm vs Kustomize compared for 2026 - templating vs patching, Helm 4's new features, and why most production teams end up running both.

5 min read•Aug 2026
Prometheus vs Datadog vs New Relic: Real Costs

Prometheus vs Datadog vs New Relic: Real Costs

Prometheus, Datadog, and New Relic compared for 2026 - real pricing at scale, hidden cost drivers, and which fits a Kubernetes-heavy stack.

5 min read•Aug 2026
Cluster Autoscaler vs Karpenter for EKS in 2026

Cluster Autoscaler vs Karpenter for EKS in 2026

Cluster Autoscaler vs Karpenter compared for EKS in 2026 - provisioning speed, bin-packing, cloud support, and when each is the right default.

5 min read•Aug 2026
GKE vs EKS vs AKS in 2026: Which Fits Your Team?

GKE vs EKS vs AKS in 2026: Which Fits Your Team?

GKE, EKS, and AKS compared for 2026 - control plane pricing, Autopilot vs Karpenter vs Node Auto Provisioning, and which platform actually fits your team.

5 min read•Aug 2026
K3s vs K8s vs MicroK8s in 2026

K3s vs K8s vs MicroK8s in 2026

K3s, full Kubernetes, and MicroK8s compared for 2026 - resource footprint, production readiness, and which fits edge, homelab, or cloud workloads.

5 min read•Aug 2026
Canary Deployments with Argo Rollouts & Flagger

Canary Deployments with Argo Rollouts & Flagger

Ship to 5% of users first and auto-rollback in minutes — a hands-on guide to canary deployments with Argo Rollouts and Flagger on Kubernetes.

5 min read•Jun 2026
OpenTelemetry Explained: Metrics, Logs, Traces

OpenTelemetry Explained: Metrics, Logs, Traces

OpenTelemetry unifies metrics, logs, and traces under one open standard — how it works, what it replaces, and how to instrument a service in 20 minutes.

5 min read•Jun 2026
Kubernetes Cost Optimization Without Breaking SLOs

Kubernetes Cost Optimization Without Breaking SLOs

Average Kubernetes CPU utilization across production clusters is 8%. Here is the complete 2026 playbook for cutting cloud spend without touching your SLOs.

5 min read•Jun 2026
ArgoCD vs FluxCD: GitOps for Kubernetes in 2026

ArgoCD vs FluxCD: GitOps for Kubernetes in 2026

ArgoCD and FluxCD are the two dominant GitOps engines for Kubernetes in 2026 — this breakdown tells you exactly which one to pick and why.

10 min read•Jun 2026

Explore More in Kubernetes Workload Management

All 6 Topics

Frequently Asked Questions

Is Running StatefulSets for Databases on Kubernetes free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Running StatefulSets for Databases on Kubernetes topic cover?

Deploy and manage stateful database workloads on Kubernetes using StatefulSets with stable network identities, ordered scaling, and persistent storage.