Skip to main content

Troubleshooting Kubernetes Pod OOMKilled and CrashLoopBackOff Errors

Master standard practices for Troubleshooting Kubernetes Pod OOMKilled and CrashLoopBackOff Errors.

52 Terms

Production Best Practices & Common Pitfalls

  • Always test cross-namespace connectivity after adding any NetworkPolicy. A policy in the target namespace affects all inbound traffic including from other namespaces that previously worked without any policy.
  • Use Cilium with Hubble in production — the real-time policy trace and drop visibility is worth the migration cost. Debugging NetworkPolicies without Hubble is guesswork.
  • Label your namespaces explicitly with kubernetes.io/metadata.name — this label is auto-applied in Kubernetes 1.21+ and is required for reliable namespace-based NetworkPolicy selectors.
  • Monitor CoreDNS with Prometheus and alert on coredns_dns_response_rcode_count_total{rcode="SERVFAIL"} — a spike indicates upstream DNS failure that will cause cascading service discovery failures across the entire cluster.
  • Never run tcpdump directly on a node in production without approval — packet capture on a financial services cluster is a compliance event that must be logged and justified.
Common Mistake

Checking pod logs to diagnose networking issues. Application logs say "connection refused" or "timeout" — they cannot tell you whether the failure is at Layer 2 (CNI), Layer 3 (kube-proxy), Layer 4 (DNS), or Layer 5 (NetworkPolicy). Always use network-level tools like netshoot, not application logs, for network debugging.

Quick Reference & Troubleshooting Commands

Command Purpose
kubectl run netshoot --image=nicolaka/netshoot -it --rm -n <ns> -- bash Launch network debug container
kubectl debug -it --image=nicolaka/netshoot --target=<pod> <pod> -n <ns> -- bash Debug sharing pod's network namespace
kubectl get endpoints <service> -n <ns> Verify pods are registered as Service backends
kubectl get networkpolicies -n <ns> List all NetworkPolicies in a namespace
kubectl get pods -n kube-system | grep coredns Check CoreDNS pod health
kubectl logs -n kube-system -l k8s-app=kube-dns CoreDNS logs for DNS failure diagnosis
nslookup <service>.<ns>.svc.cluster.local Test DNS resolution from inside a pod
curl -v http://<clusterip>:<port>/path Test Service IP directly bypassing DNS
kubectl get pods -n production -o wide Show pod IPs and which node they are on
kubectl logs -n kube-system kube-proxy-<id> kube-proxy logs for Service routing issues

Resources

Velero vs etcd Snapshot: Not a Real Choice

Velero vs etcd Snapshot: Not a Real Choice

Velero and etcd snapshots protect different layers of a cluster, not the same thing. Here's what each covers and why production DR needs both.

5 min read•Aug 2026
Istio Ambient vs Linkerd in 2026

Istio Ambient vs Linkerd in 2026

Istio Ambient killed the sidecar-tax argument. The real 2026 decision is waypoint topology and Buoyant's licensing shift, not features vs simplicity.

5 min read•Aug 2026
Nginx Ingress vs Traefik vs Gateway API in 2026

Nginx Ingress vs Traefik vs Gateway API in 2026

Ingress-nginx retired in March 2026. Here's how Traefik and the Gateway API actually compare as replacements — and why "just swap it" is the wrong frame.

5 min read•Aug 2026
OPA Gatekeeper vs Kyverno: Policy Engine in 2026

OPA Gatekeeper vs Kyverno: Policy Engine in 2026

OPA Gatekeeper vs Kyverno compared for 2026 - Rego vs YAML, mutation maturity, operational overhead, and which policy engine fits your cluster.

5 min read•Aug 2026
Helm vs Kustomize in 2026: Templating vs Patching

Helm vs Kustomize in 2026: Templating vs Patching

Helm vs Kustomize compared for 2026 - templating vs patching, Helm 4's new features, and why most production teams end up running both.

5 min read•Aug 2026
Prometheus vs Datadog vs New Relic: Real Costs

Prometheus vs Datadog vs New Relic: Real Costs

Prometheus, Datadog, and New Relic compared for 2026 - real pricing at scale, hidden cost drivers, and which fits a Kubernetes-heavy stack.

5 min read•Aug 2026
Cluster Autoscaler vs Karpenter for EKS in 2026

Cluster Autoscaler vs Karpenter for EKS in 2026

Cluster Autoscaler vs Karpenter compared for EKS in 2026 - provisioning speed, bin-packing, cloud support, and when each is the right default.

5 min read•Aug 2026
GKE vs EKS vs AKS in 2026: Which Fits Your Team?

GKE vs EKS vs AKS in 2026: Which Fits Your Team?

GKE, EKS, and AKS compared for 2026 - control plane pricing, Autopilot vs Karpenter vs Node Auto Provisioning, and which platform actually fits your team.

5 min read•Aug 2026
K3s vs K8s vs MicroK8s in 2026

K3s vs K8s vs MicroK8s in 2026

K3s, full Kubernetes, and MicroK8s compared for 2026 - resource footprint, production readiness, and which fits edge, homelab, or cloud workloads.

5 min read•Aug 2026
Canary Deployments with Argo Rollouts & Flagger

Canary Deployments with Argo Rollouts & Flagger

Ship to 5% of users first and auto-rollback in minutes — a hands-on guide to canary deployments with Argo Rollouts and Flagger on Kubernetes.

5 min read•Jun 2026
OpenTelemetry Explained: Metrics, Logs, Traces

OpenTelemetry Explained: Metrics, Logs, Traces

OpenTelemetry unifies metrics, logs, and traces under one open standard — how it works, what it replaces, and how to instrument a service in 20 minutes.

5 min read•Jun 2026
Kubernetes Cost Optimization Without Breaking SLOs

Kubernetes Cost Optimization Without Breaking SLOs

Average Kubernetes CPU utilization across production clusters is 8%. Here is the complete 2026 playbook for cutting cloud spend without touching your SLOs.

5 min read•Jun 2026
ArgoCD vs FluxCD: GitOps for Kubernetes in 2026

ArgoCD vs FluxCD: GitOps for Kubernetes in 2026

ArgoCD and FluxCD are the two dominant GitOps engines for Kubernetes in 2026 — this breakdown tells you exactly which one to pick and why.

10 min read•Jun 2026

Explore More in Kubernetes Observability and Scaling

All 3 Topics

Frequently Asked Questions

Is Troubleshooting Kubernetes Pod OOMKilled and CrashLoopBackOff Errors free to learn on DevOps Network?

Yes - this topic, like everything on DevOps Network, is 100% free with no paywall or sign-up gate.

What does the Troubleshooting Kubernetes Pod OOMKilled and CrashLoopBackOff Errors topic cover?

Master standard practices for Troubleshooting Kubernetes Pod OOMKilled and CrashLoopBackOff Errors.