Skip to main content

OOMKilled

A pod termination status in Kubernetes that occurs when a container exceeds its configured memory limit, causing the Linux kernel to forcefully terminate the process. Exit code is always 137.

What OOMKilled means

OOM stands for Out Of Memory. When a container tries to use more memory than its configured limit, the Linux kernel kills the process immediately. There is no warning and no graceful shutdown, so the application gets no chance to save state or close connections.

Kubernetes does not make this decision itself. The limit is enforced by the kernel's cgroup memory controller, and Kubernetes reports the result as OOMKilled.

%%{init: {'themeVariables': {'fontSize': '18px'}, 'flowchart': {'nodeSpacing': 40, 'rankSpacing': 45, 'padding': 14}}}%% flowchart TD A["Container limit
memory: 512Mi"] --> B["Usage reaches 512Mi"] B --> C["Kernel sends SIGKILL"] C --> D["Pod status: OOMKilled
Exit code 137"] D --> E["restartPolicy: Always
Container is restarted"] D --> F["restartPolicy: Never
Container stays stopped"] class A,B,C,E,F base class D key classDef base fill:#f8fafc,stroke:#94a3b8,stroke-width:1.5px,color:#0f172a classDef key fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#0f172a

Why the exit code is 137

When a process is killed by a signal, its exit code is 128 plus the signal number. SIGKILL is signal 9, so the exit code is 137. Unlike SIGTERM, SIGKILL cannot be caught, blocked, or ignored by the process.

Signal Number Exit code Meaning
SIGTERM 15 143 Graceful shutdown requested; the app can handle it
SIGKILL 9 137 Hard kill by the kernel, as in an OOMKill or kill -9
SIGSEGV 11 139 Segmentation fault caused by a bad memory access

How to confirm an OOMKill

Bash
# Look for "OOMKilled" under Last State
kubectl describe pod api-server-7d9f8b-xkp2q -n production
# Last State: Terminated
# Reason: OOMKilled
# Exit Code: 137
# Read the reason directly
kubectl get pod api-server-7d9f8b-xkp2q -n production \
-o jsonpath='{.status.containerStatuses[0].lastState.terminated.reason}'

Setting memory limits

The requests value is what the scheduler reserves for the pod on a node. The limits value is the hard ceiling that triggers the kill.

YAML
containers:
- name: api-server
image: registry.example.com/api-server:v2.4.1
resources:
requests:
memory: "256Mi" # Reserved on the node
cpu: "200m"
limits:
memory: "512Mi" # Exceeding this = OOMKilled
cpu: "1000m"

Measure real usage before choosing a number, then add headroom:

Bash
kubectl top pods -n production --sort-by=memory
# Rule of thumb: limit = peak observed usage + 30%
# Example: peak 380Mi x 1.3 = ~500Mi, round up to 512Mi

Runtimes that manage their own heap need extra care, because the heap can grow past the container limit before the runtime reacts:

  • Node.js: set NODE_OPTIONS=--max-old-space-size=400 for a 512Mi limit (about 75-80% of the limit), so V8 garbage-collects before the kernel steps in.
  • Java: use -XX:MaxRAMPercentage=75.0 rather than a fixed -Xmx. Older JVMs that are not container-aware size the heap from the node's total RAM.

Troubleshooting

Symptom Likely cause Fix
OOMKilled soon after start Limit set too low Measure peak with kubectl top pod, add 30% headroom
Memory climbs steadily until killed Memory leak in the application Profile the app; raising the limit only delays the kill
OOMKilled only during traffic spikes Limit too tight for peak load Raise the limit or scale out with an HPA
Java or Node app killed despite a large limit Runtime heap not bounded to the container Set the heap flags above
Common Mistake

Setting the limit equal to the request, or sizing it from average usage. Any spike or GC pause above the limit kills the pod instantly, so leave 30-50% headroom over typical usage.

Remember

A pod in CrashLoopBackOff with exit code 137 is usually being OOMKilled on every restart. Fix the memory limit or the leak rather than waiting for restarts to settle.

Frequently Asked Questions

What exactly triggers OOMKilled, and why is the exit code always 137?

When a container's memory usage hits its configured `resources.limits.memory`, the Linux kernel's cgroup memory controller, not Kubernetes itself, sends SIGKILL to the process to enforce the limit. Exit code 137 is a fixed convention (128 plus signal number 9 for SIGKILL), so it is always 137 regardless of which application or language triggered it; it is an OS-level enforcement mechanism, not an application-level exception.

What's a common mistake that causes repeated OOMKilled events in production?

Setting a memory limit based on average observed usage rather than peak usage is the most frequent cause. A JVM app under GC pressure, or a batch job processing a large payload, can briefly spike well above its steady-state footprint and get killed even though it normally fits fine. A related mistake is setting `requests` and `limits` to the same value for JVM-based apps without also tuning the JVM's own heap flags, since the JVM can attempt to use more native memory than the container limit allows. ===