OOMKilled
A pod termination status in Kubernetes that occurs when a container exceeds its configured memory limit, causing the Linux kernel to forcefully terminate the process. Exit code is always 137.
What OOMKilled means
OOM stands for Out Of Memory. When a container tries to use more memory than its configured limit, the Linux kernel kills the process immediately. There is no warning and no graceful shutdown, so the application gets no chance to save state or close connections.
Kubernetes does not make this decision itself. The limit is enforced by the kernel's cgroup memory controller, and Kubernetes reports the result as OOMKilled.
memory: 512Mi"] --> B["Usage reaches 512Mi"] B --> C["Kernel sends SIGKILL"] C --> D["Pod status: OOMKilled
Exit code 137"] D --> E["restartPolicy: Always
Container is restarted"] D --> F["restartPolicy: Never
Container stays stopped"] class A,B,C,E,F base class D key classDef base fill:#f8fafc,stroke:#94a3b8,stroke-width:1.5px,color:#0f172a classDef key fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#0f172a
Why the exit code is 137
When a process is killed by a signal, its exit code is 128 plus the signal number. SIGKILL is signal 9, so the exit code is 137. Unlike SIGTERM, SIGKILL cannot be caught, blocked, or ignored by the process.
| Signal | Number | Exit code | Meaning |
|---|---|---|---|
| SIGTERM | 15 | 143 | Graceful shutdown requested; the app can handle it |
| SIGKILL | 9 | 137 | Hard kill by the kernel, as in an OOMKill or kill -9 |
| SIGSEGV | 11 | 139 | Segmentation fault caused by a bad memory access |
How to confirm an OOMKill
# Look for "OOMKilled" under Last Statekubectl describe pod api-server-7d9f8b-xkp2q -n production # Last State: Terminated# Reason: OOMKilled# Exit Code: 137 # Read the reason directlykubectl get pod api-server-7d9f8b-xkp2q -n production \ -o jsonpath='{.status.containerStatuses[0].lastState.terminated.reason}'Setting memory limits
The requests value is what the scheduler reserves for the pod on a node. The limits value is the hard ceiling that triggers the kill.
containers: - name: api-server image: registry.example.com/api-server:v2.4.1 resources: requests: memory: "256Mi" # Reserved on the node cpu: "200m" limits: memory: "512Mi" # Exceeding this = OOMKilled cpu: "1000m"Measure real usage before choosing a number, then add headroom:
kubectl top pods -n production --sort-by=memory# Rule of thumb: limit = peak observed usage + 30%# Example: peak 380Mi x 1.3 = ~500Mi, round up to 512MiRuntimes that manage their own heap need extra care, because the heap can grow past the container limit before the runtime reacts:
- Node.js: set
NODE_OPTIONS=--max-old-space-size=400for a 512Mi limit (about 75-80% of the limit), so V8 garbage-collects before the kernel steps in. - Java: use
-XX:MaxRAMPercentage=75.0rather than a fixed-Xmx. Older JVMs that are not container-aware size the heap from the node's total RAM.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| OOMKilled soon after start | Limit set too low | Measure peak with kubectl top pod, add 30% headroom |
| Memory climbs steadily until killed | Memory leak in the application | Profile the app; raising the limit only delays the kill |
| OOMKilled only during traffic spikes | Limit too tight for peak load | Raise the limit or scale out with an HPA |
| Java or Node app killed despite a large limit | Runtime heap not bounded to the container | Set the heap flags above |
Common MistakeSetting the limit equal to the request, or sizing it from average usage. Any spike or GC pause above the limit kills the pod instantly, so leave 30-50% headroom over typical usage.
RememberA pod in CrashLoopBackOff with exit code 137 is usually being OOMKilled on every restart. Fix the memory limit or the leak rather than waiting for restarts to settle.
Frequently Asked Questions
What exactly triggers OOMKilled, and why is the exit code always 137?
When a container's memory usage hits its configured `resources.limits.memory`, the Linux kernel's cgroup memory controller, not Kubernetes itself, sends SIGKILL to the process to enforce the limit. Exit code 137 is a fixed convention (128 plus signal number 9 for SIGKILL), so it is always 137 regardless of which application or language triggered it; it is an OS-level enforcement mechanism, not an application-level exception.
What's a common mistake that causes repeated OOMKilled events in production?
Setting a memory limit based on average observed usage rather than peak usage is the most frequent cause. A JVM app under GC pressure, or a batch job processing a large payload, can briefly spike well above its steady-state footprint and get killed even though it normally fits fine. A related mistake is setting `requests` and `limits` to the same value for JVM-based apps without also tuning the JVM's own heap flags, since the JVM can attempt to use more native memory than the container limit allows. ===