Linux / shell
OOMKilled and exit code 137 in Kubernetes
Written and reviewed by Sahil Srivastav
Reason: OOMKilled
Exit Code: 137What this error actually means
The container runtime recorded an out-of-memory termination. Linux enforces container memory through control groups; when memory cannot be reclaimed within the applicable budget, the kernel can kill a process. Kubernetes records the terminated container state and, when its restart policy requires it, starts another instance. The new process has a clean memory baseline, which can hide the peak that killed its predecessor.
Exit code 137 alone is weaker evidence. It conventionally represents termination by signal 9, SIGKILL, but can also be an explicit exit status. A forced shutdown after the termination grace period, an operator’s kill, and an OOM kill can all end with 137. Read the termination reason, events and kernel evidence together before calling every such exit a memory leak.
A container limit and a node’s available memory are separate boundaries. A pod can hit its own limit while the host has gigabytes free, or become a victim of host memory pressure without reaching its configured limit. Increasing a request influences scheduling; it does not automatically raise an existing limit or repair allocations that grow without bound.
Causes, most common first
- 1Working set or concurrency exceeds the container budget. One request fits, but many concurrent requests retain buffers and query results simultaneously. A new batch size or traffic pattern changes peak resident memory without changing the steady-state graph. Rank this ahead of a leak when deaths align with specific jobs or bursts.
- 2Retained objects or queues grow with process age. Caches without eviction, queued retries and abandoned request objects accumulate. A rising baseline after idle periods distinguishes retention from temporary allocation. Increasing the limit postpones the same incident while increasing the eventual blast radius.
- 3Memory outside the managed heap was omitted. Native buffers, thread stacks, subprocesses and memory-backed emptyDir volumes can consume the same constrained budget. A heap dashboard is only one component of container memory; a flat managed heap cannot clear the rest of the process.
- 4The node ran out of usable memory. Overcommitted workloads or host processes can cause a system-level OOM event. Inspect the node before changing one workload’s limit. Kubelet eviction and kernel OOM killing are different paths and leave different evidence.
When you see it
- The previous container state says OOMKilled while the current container is Running
- Logs stop abruptly without the application’s normal shutdown sequence
- Memory looks low immediately after each restart and climbs with accumulated work
- A large import or several concurrent requests trigger kills between monitoring samples
How to diagnose it
Step 1
Preserve the previous termination record
Replace demo, app-pod and app with the affected namespace, pod and container. Read each container separately, including sidecars. Capture this before deleting the pod, because a replacement pod does not inherit its predecessor’s status.
kubectl -n demo get pod app-pod -o jsonpath='{range .status.containerStatuses[*]}{.name}{"\t"}{.lastState.terminated.reason}{"\t"}{.lastState.terminated.exitCode}{"\n"}{end}'
kubectl -n demo logs app-pod -c app --previous --timestampsStep 2
Compare the admitted resource budget with events
Inspect the live object rather than only the deployment source: namespace defaults and admission policies may have changed its resources. Correlate the termination timestamp with probes, shutdown events and node conditions.
kubectl -n demo describe pod app-pod
kubectl -n demo get pod app-pod -o jsonpath='{.spec.containers[*].resources}{"\n"}'Step 3
Read cgroup evidence where it is exposed
On a Linux cgroup v2 container with its cgroup mounted at this path, memory.events shows OOM counters and memory.max shows the enforced ceiling. A restarted container can have new counters; zero now does not disprove a previous kill. Use node-side retained telemetry for historical attribution.
kubectl -n demo exec app-pod -c app -- cat /sys/fs/cgroup/memory.events /sys/fs/cgroup/memory.current /sys/fs/cgroup/memory.maxStep 4
Capture peaks and classify the allocation shape
Graph container usage, concurrency and request sizes on the same timeline. kubectl top is a sampled snapshot, not a peak recorder. Reproduce the suspect workload with bounded concurrency and profile retained objects or native allocation before the kill occurs.
kubectl -n demo top pod app-pod --containersThe fix
Bound the multiplication that creates the peak: paginate results, stream uploads, cap worker concurrency and give queues finite capacity with an explicit overload response. Verify the total memory of simultaneous operations rather than measuring one successful request in isolation.
For retention, identify what still owns the objects after work completes and repair that lifecycle. For memory-backed files, bound and clean their contents. Moving a cache into a tmpfs volume moves the allocation mechanism without creating a new memory budget.
Raise a limit only when profiling establishes a legitimate working set and the node has capacity. Set a realistic request as well so the scheduler can place the workload honestly. Load-test the replacement under the actual limit; testing on an unrestricted workstation cannot validate the budget.
If 137 came from forced termination, repair signal delivery or draining instead. A larger memory limit cannot make an application observe SIGTERM, and it obscures the real correlation with deployment or scale-down events.
How to stop it coming back
- Alert on restarts with termination reason and container name, not only pod availability
- Record memory peaks and workload concurrency at a resolution that catches batch spikes
- Exercise the maximum supported request size under the production memory limit
- Keep memory requests, limits and runtime heap settings in the same capacity review
FAQ
Does exit 137 prove an OOM kill?
No. It commonly indicates SIGKILL, which has several causes. OOMKilled plus matching cgroup or kernel evidence establishes memory pressure much more strongly than the exit code alone.
Why is there no heap dump?
SIGKILL gives the runtime no opportunity to execute handlers or create a dump. Heap-dump-on-OutOfMemoryError only helps when the JVM itself throws that exception; the kernel can kill it before that happens.
Why did increasing the memory request not help?
A request reserves scheduling capacity and affects some pressure decisions. An explicitly configured container memory limit remains the same enforcement boundary until you change it.
Related
Other errors engineers hit next to this one
- AssertionError: daemonic processes are not allowed to have children
- [CRITICAL] WORKER TIMEOUT (pid:1234)
- Mutable default argument retains state across calls
- CommitFailedException: Commit cannot be completed since the group has already rebalanced
- Consumer group stuck rebalancing — poll timeout has expired
- The same message processed twice (at-least-once delivery)
- Messages processed out of order across partitions
- Webhook delivered twice — customer charged twice