Skip to main content

Configuring Persistent Volumes and Storage Classes in Kubernetes

Configure PersistentVolumes, PersistentVolumeClaims, and StorageClasses in Kubernetes to provide durable storage for stateful workloads in production.

52 Terms

Overview and What You Will Learn

Containers are ephemeral by design — when a pod restarts, all data written to its filesystem is lost. For databases, message queues, and any stateful workload, this is catastrophic. Kubernetes solves this through PersistentVolumes (PV), PersistentVolumeClaims (PVC), and StorageClasses — a three-layer abstraction that decouples storage provisioning from storage consumption, allowing pods to survive restarts, rescheduling, and node failures without losing data.

By the end of this guide you will be able to:

  • Understand the PV, PVC, and StorageClass relationship and provisioning lifecycle
  • Create StorageClasses for dynamic volume provisioning on AWS, GCP, and on-prem clusters
  • Write PersistentVolumeClaims and mount volumes correctly inside pod specs
  • Configure volume access modes and reclaim policies for different production workloads
  • Troubleshoot PVC stuck in Pending state and volume mount failures

Why This Matters in Production

Zerodha's trading platform stores order books, trade history, and user portfolio data in PostgreSQL running on Kubernetes. If the PostgreSQL pod restarts without a PersistentVolume, every trade record since the last external backup is lost — a regulatory violation and a catastrophic user trust failure.

At Hotstar, the video transcoding pipeline writes intermediate encoded segments to shared storage that multiple pods must read simultaneously. The wrong access mode on the PVC causes silent data corruption or outright mount failures. Understanding storage configuration is not optional for any engineer running stateful workloads on Kubernetes.

Core Principles

The three-layer storage abstraction and how they compose:

◈ DIAGRAM
CLUSTER ADMIN
StorageClass
(defines HOW storage is provisioned:
AWS EBS, GCP PD, NFS, local disk)
│
▼
DEVELOPER
PersistentVolumeClaim
(requests WHAT storage is needed:
size, access mode, storage class)
│
▼
APPLICATION
Pod spec
(mounts the PVC as a volume at a path)
│
▼
PersistentVolume (PV)
(the actual provisioned storage unit —
created automatically by StorageClass,
or manually by admin for static provisioning)

Access modes — the most misunderstood configuration in Kubernetes storage:

ReadWriteOnce (RWO) → One node can mount read-write at a time Used for: databases, single-instance stateful apps Supported by: AWS EBS, GCP Persistent Disk, Azure Disk

ReadWriteMany (RWX) → Multiple nodes can mount read-write simultaneously Used for: shared file storage, media assets, ML datasets Supported by: NFS, AWS EFS, GCP Filestore, Azure Files NOT supported by: EBS, GCP PD, Azure Disk

ReadOnlyMany (ROX) → Multiple nodes can mount read-only simultaneously Used for: shared config files, static assets

Detailed Step-by-Step Practical Lab

Step 1 — Inspect Available StorageClasses
Bash
kubectl get storageclasses

Example output on an AWS EKS cluster:

Bash
NAME PROVISIONER RECLAIMPOLICY VOLUMEBINDINGMODE
gp2 (default) ebs.csi.aws.com Delete WaitForFirstConsumer
gp3-encrypted ebs.csi.aws.com Retain WaitForFirstConsumer
efs-sc efs.csi.aws.com Retain Immediate

Inspect a specific StorageClass for full configuration:

Bash
kubectl describe storageclass gp3-encrypted
Remember

RECLAIMPOLICY: Delete means the underlying cloud disk is permanently deleted when the PVC is deleted. RECLAIMPOLICY: Retain keeps the disk even after PVC deletion — always use Retain for production databases.

Step 2 — Create Production StorageClasses
YAML
# storageclasses.yaml — define storage tiers for different workload types
# Tier 1: Fast encrypted SSD for databases (AWS gp3)
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: gp3-encrypted
annotations:
storageclass.kubernetes.io/is-default-class: "false"
provisioner: ebs.csi.aws.com
parameters:
type: gp3
iops: "6000" # 6000 IOPS — good for PostgreSQL/MySQL
throughput: "250" # 250 MB/s throughput
encrypted: "true" # Encrypt at rest — required for financial data
kmsKeyId: "arn:aws:kms:ap-south-1:123456789:key/zerodha-ebs-key"
reclaimPolicy: Retain # NEVER auto-delete production database disks
allowVolumeExpansion: true # Allow resizing PVCs without downtime
volumeBindingMode: WaitForFirstConsumer # Provision in same AZ as pod
---
# Tier 2: Shared file storage for media assets (AWS EFS — supports RWX)
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: efs-shared
provisioner: efs.csi.aws.com
parameters:
provisioningMode: efs-ap
fileSystemId: fs-0a1b2c3d4e5f6789 # Your EFS filesystem ID
directoryPerms: "700"
reclaimPolicy: Retain
volumeBindingMode: Immediate
---
# Tier 3: Fast local NVMe for temporary high-performance workloads
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: local-nvme
provisioner: kubernetes.io/no-provisioner # Manual provisioning
volumeBindingMode: WaitForFirstConsumer
reclaimPolicy: Delete # Local storage is node-specific — delete on release
Bash
kubectl apply -f storageclasses.yaml
Step 3 — Create PersistentVolumeClaims for Different Workloads
YAML
# pvcs.yaml — storage claims for different production workloads
# PVC for PostgreSQL database — single node, high IOPS, encrypted
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: postgres-data-pvc
namespace: production
labels:
app: postgres
team: platform
spec:
accessModes:
- ReadWriteOnce # Only one node mounts at a time — correct for databases
storageClassName: gp3-encrypted
resources:
requests:
storage: 100Gi # Start with 100GB — can expand later without downtime
---
# PVC for Hotstar video asset storage — multiple pods read/write simultaneously
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: video-assets-pvc
namespace: transcoding
spec:
accessModes:
- ReadWriteMany # Multiple transcoding pods mount simultaneously
storageClassName: efs-shared
resources:
requests:
storage: 5Ti # 5TB for video asset storage
---
# PVC for Redis cache persistence — small, fast, single node
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: redis-data-pvc
namespace: production
spec:
accessModes:
- ReadWriteOnce
storageClassName: gp3-encrypted
resources:
requests:
storage: 20Gi
Bash
kubectl apply -f pvcs.yaml

Check PVC status — should move from Pending to Bound:

Bash
kubectl get pvc -n production
TEXT
NAME STATUS VOLUME CAPACITY
postgres-data-pvc Bound pvc-a1b2c3d4-e5f6-7890-abcd-ef1234567890 100Gi
redis-data-pvc Bound pvc-b2c3d4e5-f6a7-8901-bcde-f12345678901 20Gi
Step 4 — Mount PVCs into Pod Specs
YAML
# deployment-postgres.yaml — PostgreSQL with persistent storage
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: postgres
namespace: production
spec:
serviceName: postgres
replicas: 1
selector:
matchLabels:
app: postgres
template:
metadata:
labels:
app: postgres
spec:
securityContext:
fsGroup: 999 # PostgreSQL runs as UID 999 — set volume ownership
containers:
- name: postgres
image: postgres:15.4
env:
- name: POSTGRES_DB
value: zerodha_trading
- name: POSTGRES_USER
value: rahul
- name: POSTGRES_PASSWORD
valueFrom:
secretKeyRef:
name: postgres-credentials
key: password
- name: PGDATA
value: /var/lib/postgresql/data/pgdata # Subdirectory avoids lost+found issue
ports:
- containerPort: 5432
resources:
requests:
cpu: "500m"
memory: "1Gi"
limits:
cpu: "2"
memory: "4Gi"
volumeMounts:
- name: postgres-storage
mountPath: /var/lib/postgresql/data # PostgreSQL data directory
volumes:
- name: postgres-storage
persistentVolumeClaim:
claimName: postgres-data-pvc # Reference the PVC by name
Bash
kubectl apply -f deployment-postgres.yaml

Verify volume is mounted inside the pod:

Bash
kubectl exec -it postgres-0 -n production -- df -h /var/lib/postgresql/data
TEXT
Filesystem Size Used Avail Use% Mounted on
/dev/nvme1n1 98G 156M 98G 1% /var/lib/postgresql/data
Security

Always set securityContext.fsGroup to match the UID your application runs as. Without it, the mounted volume is owned by root and your application process may fail to write to it — causing a crash that looks like a storage failure but is actually a permissions issue.

Step 5 — Expand a PVC Without Downtime

When your database grows beyond the initial allocation:

Verify the StorageClass supports volume expansion:

Bash
kubectl get storageclass gp3-encrypted -o jsonpath='{.allowVolumeExpansion}'
TEXT
true

Edit the PVC to request more storage:

Bash
kubectl patch pvc postgres-data-pvc -n production \
--type='json' \
-p='[{"op":"replace","path":"/spec/resources/requests/storage","value":"200Gi"}]'

Watch the expansion happen:

Bash
kubectl get pvc postgres-data-pvc -n production -w
◈ DIAGRAM
NAME STATUS CAPACITY CONDITIONS
postgres-data-pvc Bound 100Gi Resizing...
postgres-data-pvc Bound 200Gi FileSystemResizePending
postgres-data-pvc Bound 200Gi ← expansion complete

For filesystem resize to complete — the pod may need a restart:

Bash
kubectl rollout restart statefulset/postgres -n production
Tip

Volume expansion only works in one direction — you can increase a PVC's size but never decrease it. Always start with a reasonable baseline and use allowVolumeExpansion: true on your StorageClass so you can grow without recreating the PVC.

Step 6 — Troubleshoot PVC Stuck in Pending State

PVC not moving from Pending to Bound:

Bash
kubectl get pvc postgres-data-pvc -n production
◈ DIAGRAM
NAME STATUS VOLUME CAPACITY ACCESS MODES
postgres-data-pvc Pending ← stuck

Step 1 — Describe the PVC for the reason:

Bash
kubectl describe pvc postgres-data-pvc -n production
TEXT
Events:
Warning ProvisioningFailed storageclass.storage.k8s.io "gp3-encrypted" not found

→ StorageClass name is wrong or not installed

TEXT
Warning ProvisioningFailed failed to provision volume:
InvalidParameterValue: The iops parameter is not supported for volume type gp2

→ Wrong parameters for the volume type

TEXT
Warning WaitForFirstConsumer waiting for first consumer to be created

→ VolumeBindingMode is WaitForFirstConsumer — PVC will stay Pending until a pod tries to mount it. This is normal.

Step 2 — Check if the CSI driver is running:

Bash
kubectl get pods -n kube-system | grep ebs-csi
TEXT
ebs-csi-controller-xxx 6/6 Running 0 5d
ebs-csi-node-xxx 3/3 Running 0 5d

Step 3 — Check CSI driver logs for provisioning errors:

Bash
kubectl logs -n kube-system \
-l app=ebs-csi-controller \
-c csi-provisioner \
--tail=50
Step 7 — Implement Volume Snapshots for Backup
YAML
# volume-snapshot.yaml — take a point-in-time snapshot of the PostgreSQL volume
apiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
name: postgres-snapshot-20250525
namespace: production
spec:
volumeSnapshotClassName: csi-aws-vsc
source:
persistentVolumeClaimName: postgres-data-pvc # Snapshot this PVC
Bash
kubectl apply -f volume-snapshot.yaml

Check snapshot status:

Bash
kubectl get volumesnapshot -n production
TEXT
NAME READYTOUSE SOURCEPVC RESTORESIZE AGE
postgres-snapshot-20250525 true postgres-data-pvc 100Gi 2m

Restore from snapshot into a new PVC:

Bash
kubectl apply -f - <<EOF
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: postgres-data-restored
namespace: production
spec:
accessModes:
- ReadWriteOnce
storageClassName: gp3-encrypted
resources:
requests:
storage: 100Gi
dataSource:
name: postgres-snapshot-20250525
kind: VolumeSnapshot
apiGroup: snapshot.storage.k8s.io
EOF
Remember

Volume snapshots are crash-consistent, not application-consistent. For PostgreSQL, always run pg_dump or use pg_basebackup for application-consistent backups. Use volume snapshots as a fast recovery complement, not as your only backup strategy.

Production Best Practices & Common Pitfalls

  • Always use PGDATA=/var/lib/postgresql/data/pgdata (a subdirectory) for PostgreSQL on Kubernetes. Mounting directly to /var/lib/postgresql/data causes PostgreSQL to fail because the volume root contains a lost+found directory it cannot handle.
  • Tag your PVCs with team and application labels — at scale, identifying which PVC belongs to which application becomes impossible without consistent labelling.
  • Set up automated volume snapshot schedules using Velero or the cloud provider's native snapshot scheduler. A 100GB PostgreSQL volume with no snapshots is a single point of failure.
  • Monitor PVC usage with kubectl exec <pod> -- df -h and alert at 80% full — Kubernetes does not automatically expand volumes and a full disk causes immediate pod failure.
  • Never share a single RWO PVC between multiple pods. Only one node can mount it at a time — if a second pod tries to mount it on a different node, it will stay in Pending or ContainerCreating indefinitely.
Common Mistake

Using reclaimPolicy: Delete on production database StorageClasses. When a developer accidentally runs kubectl delete pvc postgres-data-pvc, the underlying cloud disk and all its data is permanently deleted within seconds. Always use Retain for any storage containing production data.

Quick Reference & Troubleshooting Commands

Command Purpose
kubectl get pvc -n <ns> List all PVCs and their binding status
kubectl describe pvc <name> -n <ns> Full PVC details and provisioning events
kubectl get pv List all PersistentVolumes cluster-wide
kubectl get storageclass List available StorageClasses
kubectl describe storageclass <name> Full StorageClass configuration
kubectl patch pvc <name> -n <ns> --type='json' -p='[...]' Expand PVC size
kubectl exec <pod> -n <ns> -- df -h Check disk usage inside a pod
kubectl get volumesnapshot -n <ns> List volume snapshots
kubectl logs -n kube-system -l app=ebs-csi-controller -c csi-provisioner Debug CSI provisioning failures
kubectl get events -n <ns> --field-selector reason=ProvisioningFailed Filter storage provisioning failures

Resources

Velero vs etcd Snapshot: Not a Real Choice

Velero vs etcd Snapshot: Not a Real Choice

Velero and etcd snapshots protect different layers of a cluster, not the same thing. Here's what each covers and why production DR needs both.

5 min read•Aug 2026
Istio Ambient vs Linkerd in 2026

Istio Ambient vs Linkerd in 2026

Istio Ambient killed the sidecar-tax argument. The real 2026 decision is waypoint topology and Buoyant's licensing shift, not features vs simplicity.

5 min read•Aug 2026
Nginx Ingress vs Traefik vs Gateway API in 2026

Nginx Ingress vs Traefik vs Gateway API in 2026

Ingress-nginx retired in March 2026. Here's how Traefik and the Gateway API actually compare as replacements — and why "just swap it" is the wrong frame.

5 min read•Aug 2026
OPA Gatekeeper vs Kyverno: Policy Engine in 2026

OPA Gatekeeper vs Kyverno: Policy Engine in 2026

OPA Gatekeeper vs Kyverno compared for 2026 - Rego vs YAML, mutation maturity, operational overhead, and which policy engine fits your cluster.

5 min read•Aug 2026
Helm vs Kustomize in 2026: Templating vs Patching

Helm vs Kustomize in 2026: Templating vs Patching

Helm vs Kustomize compared for 2026 - templating vs patching, Helm 4's new features, and why most production teams end up running both.

5 min read•Aug 2026
Prometheus vs Datadog vs New Relic: Real Costs

Prometheus vs Datadog vs New Relic: Real Costs

Prometheus, Datadog, and New Relic compared for 2026 - real pricing at scale, hidden cost drivers, and which fits a Kubernetes-heavy stack.

5 min read•Aug 2026
Cluster Autoscaler vs Karpenter for EKS in 2026

Cluster Autoscaler vs Karpenter for EKS in 2026

Cluster Autoscaler vs Karpenter compared for EKS in 2026 - provisioning speed, bin-packing, cloud support, and when each is the right default.

5 min read•Aug 2026
GKE vs EKS vs AKS in 2026: Which Fits Your Team?

GKE vs EKS vs AKS in 2026: Which Fits Your Team?

GKE, EKS, and AKS compared for 2026 - control plane pricing, Autopilot vs Karpenter vs Node Auto Provisioning, and which platform actually fits your team.

5 min read•Aug 2026
K3s vs K8s vs MicroK8s in 2026

K3s vs K8s vs MicroK8s in 2026

K3s, full Kubernetes, and MicroK8s compared for 2026 - resource footprint, production readiness, and which fits edge, homelab, or cloud workloads.

5 min read•Aug 2026
Canary Deployments with Argo Rollouts & Flagger

Canary Deployments with Argo Rollouts & Flagger

Ship to 5% of users first and auto-rollback in minutes — a hands-on guide to canary deployments with Argo Rollouts and Flagger on Kubernetes.

5 min read•Jun 2026
OpenTelemetry Explained: Metrics, Logs, Traces

OpenTelemetry Explained: Metrics, Logs, Traces

OpenTelemetry unifies metrics, logs, and traces under one open standard — how it works, what it replaces, and how to instrument a service in 20 minutes.

5 min read•Jun 2026
Kubernetes Cost Optimization Without Breaking SLOs

Kubernetes Cost Optimization Without Breaking SLOs

Average Kubernetes CPU utilization across production clusters is 8%. Here is the complete 2026 playbook for cutting cloud spend without touching your SLOs.

5 min read•Jun 2026
ArgoCD vs FluxCD: GitOps for Kubernetes in 2026

ArgoCD vs FluxCD: GitOps for Kubernetes in 2026

ArgoCD and FluxCD are the two dominant GitOps engines for Kubernetes in 2026 — this breakdown tells you exactly which one to pick and why.

10 min read•Jun 2026

Explore More in Kubernetes Workload Management

All 6 Topics

Frequently Asked Questions

Is Configuring Persistent Volumes and Storage Classes in Kubernetes free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Configuring Persistent Volumes and Storage Classes in Kubernetes topic cover?

Configure PersistentVolumes, PersistentVolumeClaims, and StorageClasses in Kubernetes to provide durable storage for stateful workloads in production.