Skip to main content

Configuring Pod Disruption Budgets for Zero-Downtime Upgrades

A Pod Disruption Budget (PDB) is a policy that tells Kubernetes the minimum number of pods that must stay running during voluntary disruptions — node drains, cluster upgrades, and rolling deployments. Without one, a node drain can terminate all pods of a service simultaneously, causing a complete outage.

52 Terms

A Pod Disruption Budget (PDB) is a policy that tells Kubernetes the minimum number of pods that must stay running during voluntary disruptions — node drains, cluster upgrades, and rolling deployments. Without one, a node drain can terminate all pods of a service simultaneously, causing a complete outage.

+++

Configuring Pod Disruption Budgets for Zero-Downtime Upgrades

The Problem PDB Solves

Imagine your payments API has 3 pods spread across 3 nodes. A cluster upgrade requires draining all nodes one by one. Without a PDB, Kubernetes can drain Node 1, terminating its pod — that is fine, you still have 2. But it can immediately drain Node 2 next. Now you have 1 pod serving all production traffic. Then Node 3. Zero pods. Complete outage.

A PDB prevents this by telling the cluster: "Never let availability drop below 2 pods while you drain nodes."

◈ DIAGRAM
WITHOUT PDB: WITH PDB (minAvailable: 2):
Node drain sequence: Node drain sequence:
Node-1 drained → 2 pods running Node-1 drained → 2 pods running ✓
Node-2 drained → 1 pod running Node-2 drain attempt:
Node-3 drained → 0 pods ← OUTAGE Kubernetes checks PDB
2 pods available = minimum met
WAIT — cannot proceed
New pod scheduled first
3 pods running again
Node-2 drained → 2 pods ✓

Voluntary vs Involuntary Disruptions

PDBs only apply to voluntary disruptions — actions an administrator or the cluster itself initiates intentionally.

Type Examples PDB Applies?
Voluntary kubectl drain, cluster upgrade, node scaling down, admin deletes pod ✅ Yes
Involuntary Node hardware failure, kernel panic, out-of-memory kill ❌ No

PDB cannot protect you from a node dying unexpectedly. It only governs intentional operations.

Two Ways to Define a PDB

Option 1 — minAvailable: At least this many pods must be running at all times.

YAML
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: payments-api-pdb
namespace: production
spec:
minAvailable: 2 # At least 2 pods must be available during any disruption
selector:
matchLabels:
app: payments-api # Targets pods with this label

Option 2 — maxUnavailable: At most this many pods can be down at the same time.

YAML
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: payments-api-pdb
namespace: production
spec:
maxUnavailable: 1 # Only 1 pod can be unavailable at a time
selector:
matchLabels:
app: payments-api

`minAvailable` vs `maxUnavailable` — Which to Use

◈ DIAGRAM
Deployment: 5 replicas
minAvailable: 3
→ Kubernetes can disrupt at most 2 pods at a time
→ Absolute number — stays fixed even if you scale the deployment
maxUnavailable: 1
→ Kubernetes can disrupt at most 1 pod at a time
→ Percentage option: maxUnavailable: "20%" adjusts as replicas scale
Setting Best For Watch Out
minAvailable: N Critical services where you know the exact floor (e.g. "always 2 payment pods") If replicas drop below N for any reason, node drains will block indefinitely
maxUnavailable: N Services where you want proportional safety as replicas scale Less intuitive for ops teams to reason about in an incident
maxUnavailable: "10%" Large deployments (20+ replicas) Rounds down — 10% of 5 pods = 0, meaning nothing can be disrupted
Critical mistake

: Setting minAvailable equal to your replica count. Example: 3 replicas with minAvailable: 3. Kubernetes can never drain a node because draining any node would violate the budget. Cluster upgrades will stall permanently until someone deletes the PDB.

Using Percentages

YAML
spec:
minAvailable: "60%" # At least 60% of matched pods must be available
# With 10 replicas: at least 6 must be running → up to 4 can be disrupted
# With 5 replicas: at least 3 must be running → up to 2 can be disrupted
# With 3 replicas: at least 2 must be running → up to 1 can be disrupted

Percentages are useful for autoscaled deployments where replica count fluctuates — your PDB stays proportionally correct without manual updates.

Checking PDB Status

Bash
# List all PDBs in a namespace
kubectl get pdb -n production
# Output:
# NAME MIN AVAILABLE MAX UNAVAILABLE ALLOWED DISRUPTIONS AGE
# payments-api-pdb 2 N/A 1 5d
# auth-service-pdb N/A 1 1 5d
# "ALLOWED DISRUPTIONS" = how many pods can currently be taken down
# If this is 0, node drains will block
# Describe for full details
kubectl describe pdb payments-api-pdb -n production

Why a Node Drain Gets Stuck

The most common scenario at Razorpay or Hotstar: a cluster upgrade is running, and one node refuses to drain. The drain command hangs. The reason is almost always a PDB with ALLOWED DISRUPTIONS: 0.

Bash
# Drain a node during cluster upgrade
kubectl drain node mumbai-worker-3 \
--ignore-daemonsets \
--delete-emptydir-data
# Output when PDB is blocking:
# error when evicting pods/"payments-api-7d9f8b-xk2p9" -n "production"
# (will retry after 5s): Cannot evict pod as it would violate
# the pod's disruption budget.
Bash
# Diagnose why ALLOWED DISRUPTIONS is 0
kubectl get pdb payments-api-pdb -n production
# Then check the actual pod count vs minAvailable
kubectl get pods -l app=payments-api -n production
# If only 2 pods are running and minAvailable is 2:
# No pod can be evicted — evicting any one drops below the minimum
# Fix: Scale up the deployment to 3+ replicas first, then drain
kubectl scale deployment payments-api --replicas=4 -n production

Full Production Setup — Deployment + PDB Together

YAML
# deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: payments-api
namespace: production
spec:
replicas: 3
selector:
matchLabels:
app: payments-api
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1 # Rolling update can take down 1 pod at a time
maxSurge: 1 # Can temporarily create 1 extra pod during rollout
template:
metadata:
labels:
app: payments-api # ← Must match the PDB selector exactly
spec:
containers:
- name: api
image: registry.razorpay.in/payments-api:v2.5.1
# pdb.yaml — Apply this alongside the Deployment
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: payments-api-pdb
namespace: production
spec:
maxUnavailable: 1
selector:
matchLabels:
app: payments-api # ← Must match the Deployment pod labels exactly
Bash
# Apply both together
kubectl apply -f deployment.yaml -f pdb.yaml -n production
# Verify the PDB is correctly targeting pods
kubectl get pdb payments-api-pdb -n production
# ALLOWED DISRUPTIONS should be 1 if 3 pods are running and maxUnavailable is 1

PDB for StatefulSets

StatefulSets (databases, Kafka, Zookeeper) are particularly important to protect because they have no load balancer in front — each pod is individually addressable and a quorum may be required.

YAML
# zookeeper-pdb.yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: zookeeper-pdb
namespace: production
spec:
minAvailable: 2 # Zookeeper 3-node cluster: must keep 2 for quorum
selector:
matchLabels:
app: zookeeper
TEXT
3-node Zookeeper cluster requires quorum of 2 to elect a leader.
minAvailable: 2 ensures Kubernetes never drains 2 nodes simultaneously,
which would break quorum and make the cluster read-only.

Checking PDB During Cluster Upgrade (Incident Workflow)

Bash
# 1. Before draining, check all PDB statuses across the cluster
kubectl get pdb --all-namespaces
# 2. Identify any PDB with ALLOWED DISRUPTIONS = 0
kubectl get pdb --all-namespaces | grep " 0 "
# 3. For each blocking PDB, check actual pod count
kubectl get pods -n <namespace> -l <label-from-pdb-selector>
# 4. If pods are fewer than minAvailable due to earlier failures:
kubectl scale deployment <name> --replicas=<higher-count> -n <namespace>
# 5. Then retry the drain
kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data

🔴 Common Mistake: Creating a PDB but using a selector that does not match any pods. The PDB exists but protects nothing. Always verify kubectl get pdb shows a non-zero ALLOWED DISRUPTIONS value after creation — if it shows N/A or the pod count is 0, your selector is wrong.

💡 Tip: In clusters running on EKS or GKE where the cloud provider performs node upgrades automatically, PDBs are your last line of defense against upgrade-caused outages. The cloud upgrade process respects PDBs before draining nodes. At Hotstar scale, every Deployment with more than 1 replica should have a PDB — with maxUnavailable: 1 as the safe default for stateless services.

Resources

Velero vs etcd Snapshot: Not a Real Choice

Velero vs etcd Snapshot: Not a Real Choice

Velero and etcd snapshots protect different layers of a cluster, not the same thing. Here's what each covers and why production DR needs both.

5 min read•Aug 2026
Istio Ambient vs Linkerd in 2026

Istio Ambient vs Linkerd in 2026

Istio Ambient killed the sidecar-tax argument. The real 2026 decision is waypoint topology and Buoyant's licensing shift, not features vs simplicity.

5 min read•Aug 2026
Nginx Ingress vs Traefik vs Gateway API in 2026

Nginx Ingress vs Traefik vs Gateway API in 2026

Ingress-nginx retired in March 2026. Here's how Traefik and the Gateway API actually compare as replacements — and why "just swap it" is the wrong frame.

5 min read•Aug 2026
OPA Gatekeeper vs Kyverno: Policy Engine in 2026

OPA Gatekeeper vs Kyverno: Policy Engine in 2026

OPA Gatekeeper vs Kyverno compared for 2026 - Rego vs YAML, mutation maturity, operational overhead, and which policy engine fits your cluster.

5 min read•Aug 2026
Helm vs Kustomize in 2026: Templating vs Patching

Helm vs Kustomize in 2026: Templating vs Patching

Helm vs Kustomize compared for 2026 - templating vs patching, Helm 4's new features, and why most production teams end up running both.

5 min read•Aug 2026
Prometheus vs Datadog vs New Relic: Real Costs

Prometheus vs Datadog vs New Relic: Real Costs

Prometheus, Datadog, and New Relic compared for 2026 - real pricing at scale, hidden cost drivers, and which fits a Kubernetes-heavy stack.

5 min read•Aug 2026
Cluster Autoscaler vs Karpenter for EKS in 2026

Cluster Autoscaler vs Karpenter for EKS in 2026

Cluster Autoscaler vs Karpenter compared for EKS in 2026 - provisioning speed, bin-packing, cloud support, and when each is the right default.

5 min read•Aug 2026
GKE vs EKS vs AKS in 2026: Which Fits Your Team?

GKE vs EKS vs AKS in 2026: Which Fits Your Team?

GKE, EKS, and AKS compared for 2026 - control plane pricing, Autopilot vs Karpenter vs Node Auto Provisioning, and which platform actually fits your team.

5 min read•Aug 2026
K3s vs K8s vs MicroK8s in 2026

K3s vs K8s vs MicroK8s in 2026

K3s, full Kubernetes, and MicroK8s compared for 2026 - resource footprint, production readiness, and which fits edge, homelab, or cloud workloads.

5 min read•Aug 2026
Canary Deployments with Argo Rollouts & Flagger

Canary Deployments with Argo Rollouts & Flagger

Ship to 5% of users first and auto-rollback in minutes — a hands-on guide to canary deployments with Argo Rollouts and Flagger on Kubernetes.

5 min read•Jun 2026
OpenTelemetry Explained: Metrics, Logs, Traces

OpenTelemetry Explained: Metrics, Logs, Traces

OpenTelemetry unifies metrics, logs, and traces under one open standard — how it works, what it replaces, and how to instrument a service in 20 minutes.

5 min read•Jun 2026
Kubernetes Cost Optimization Without Breaking SLOs

Kubernetes Cost Optimization Without Breaking SLOs

Average Kubernetes CPU utilization across production clusters is 8%. Here is the complete 2026 playbook for cutting cloud spend without touching your SLOs.

5 min read•Jun 2026
ArgoCD vs FluxCD: GitOps for Kubernetes in 2026

ArgoCD vs FluxCD: GitOps for Kubernetes in 2026

ArgoCD and FluxCD are the two dominant GitOps engines for Kubernetes in 2026 — this breakdown tells you exactly which one to pick and why.

10 min read•Jun 2026

Explore More in Kubernetes Workload Management

All 6 Topics

Frequently Asked Questions

Is Configuring Pod Disruption Budgets for Zero-Downtime Upgrades free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Configuring Pod Disruption Budgets for Zero-Downtime Upgrades topic cover?

A Pod Disruption Budget (PDB) is a policy that tells Kubernetes the minimum number of pods that must stay running during voluntary disruptions — node drains, cluster upgrades, and rolling deployments. Without one, a node drain can terminate all pods of a service simultaneously, causing a complete outage.