CrashLoopBackOff
A Kubernetes pod status indicating the container repeatedly crashes immediately after starting and the kubelet applies exponential backoff delays before each restart attempt, beginning at 10 seconds and capping at 5 minutes. It signals that something inside the container is fundamentally broken - not a transient network issue.
CrashLoopBackOff - Why Your Pod Keeps Dying
What Does CrashLoopBackOff Mean in Simple Terms?
Your container starts, crashes, Kubernetes restarts it, it crashes again. After a few attempts Kubernetes says: "I will keep trying but I will wait longer each time so I don't waste resources." That waiting period is the backoff. The pod is not stuck — it is actively retrying on a growing timer.
Backoff Timer Progression
+-----------------------------+| Restart 1 -> wait 10s |+-----------------------------+ | v+-----------------------------+| Restart 2 -> wait 20s |+-----------------------------+ | v+-----------------------------+| Restart 3 -> wait 40s |+-----------------------------+ | v+-----------------------------+| Restart 4 -> wait 80s |+-----------------------------+ | v+-----------------------------+| Restart 5+ -> wait 300s | <- permanently capped at 5 minutes+-----------------------------+How to Diagnose CrashLoopBackOff
# Step 1 — Get logs from the PREVIOUS crashed container# The current container may have no logs yet — always use --previouskubectl logs api-server-7d9f8b-xkp2q -n production --previous # Step 2 — Check the exact exit code of the last crashkubectl get pod api-server-7d9f8b-xkp2q -n production \ -o jsonpath='{.status.containerStatuses[0].lastState.terminated.exitCode}' # Step 3 — Read the full event history for the podkubectl describe pod api-server-7d9f8b-xkp2q -n production # Step 4 — Watch restart count climb in real timekubectl get pod api-server-7d9f8b-xkp2q -n production -wExit Code Reference
| Exit Code | Meaning | Typical Fix |
|---|---|---|
1 |
Application error — exception or panic on startup | Check logs for stack trace, missing config |
2 |
Shell or script error in ENTRYPOINT | Check Dockerfile ENTRYPOINT syntax |
137 |
OOMKilled — memory limit exceeded | Increase resources.limits.memory |
139 |
Segmentation fault | Application bug — check core dump |
255 |
Entrypoint binary not found or permission denied | Verify binary path and file permissions in image |
Most Common Causes and Fixes
1. Missing environment variable — app panics on startup
# deployment.yaml — ensure all required env vars are providedspec: containers: - name: api-server image: registry.razorpay.in/api-server:v2.4.1 env: - name: DATABASE_URL valueFrom: secretKeyRef: name: db-credentials key: connection-string # Crashes if this Secret or key doesn't exist - name: REDIS_HOST value: "10.0.1.50"2. Liveness probe too aggressive — kills a slow-starting pod
# Increase initialDelaySeconds for apps that take time to bootlivenessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 30 # Give the app 30s to start before first check periodSeconds: 10 failureThreshold: 33. OOMKilled — container hitting memory limit
# Confirm OOMKill from exit codekubectl get pod <pod> -n production \ -o jsonpath='{.status.containerStatuses[0].lastState.terminated.reason}'# Output: OOMKilled # Fix: increase memory limit in deployment specresources: requests: memory: "256Mi" limits: memory: "512Mi" # Raise this if app legitimately needs more4. Image entrypoint does not exist
# Run the image locally to confirm the binary is presentdocker run --rm registry.razorpay.in/api-server:v2.4.1 ls /app/ # Override entrypoint to debug inside the containerkubectl run debug-shell \ --image=registry.razorpay.in/api-server:v2.4.1 \ --command -- sleep 3600 -n productionkubectl exec -it debug-shell -n production -- shQuick Troubleshooting Checklist
| Check | Command |
|---|---|
| Get crash logs | kubectl logs <pod> -n production --previous |
| Read exit code | kubectl get pod <pod> -o jsonpath='..lastState.terminated.exitCode' |
| Check events | kubectl describe pod <pod> -n production |
| Verify Secrets exist | kubectl get secret db-credentials -n production |
| Check resource limits | `kubectl describe pod |
TipAlways use
--previouswhen pulling logs for a CrashLoopBackOff pod. The current container restarts so fast it often has zero log output. The logs you need are always in the previous container's run.
Common MistakeRunning
kubectl logswithout--previous, seeing no output, and assuming there are no logs. This causes engineers at Swiggy and Razorpay to waste 30 minutes debugging the wrong thing. The crash logs are always there — in the previous container instance.
RememberCrashLoopBackOff is a symptom, not a root cause. Exit code
137points to memory, exit code1points to application errors, exit code255points to a broken image. Always read the exit code first — it tells you exactly which direction to debug.
SecurityIf a pod enters CrashLoopBackOff due to a missing Secret, Kubernetes will log the SecretKeyRef failure in pod events visible to anyone with
kubectl describeaccess. Avoid storing sensitive key names as environment variable names that reveal internal infrastructure topology in production namespaces.
Frequently Asked Questions
Why does CrashLoopBackOff wait longer between each restart attempt instead of retrying immediately?
Kubernetes uses exponential backoff (starting at 10 seconds, doubling up to a 5-minute cap) specifically to avoid hammering a broken container with constant restart attempts, which would waste node resources and potentially worsen the underlying problem (like overwhelming a database the container fails to connect to). The growing delay also gives whatever the container depends on — a database, a config service — more breathing room to recover if the crash is caused by that dependency being temporarily unavailable.
What's the fastest way to diagnose a CrashLoopBackOff in practice?
Check `kubectl logs <pod> --previous` first — this gets logs from the last crashed instance, not the current restarting one, and usually shows the actual error (missing environment variable, failed dependency connection, unhandled exception on startup). A common trap is assuming it's a Kubernetes networking issue when it's almost always something inside the container itself failing immediately — a missing config value, wrong file permissions, or an unhandled startup exception — not a transient infrastructure problem.