Linux / shell

OOMKilled and exit code 137 in Kubernetes

Written and reviewed by Sahil Srivastav

KubernetesMemory limitsSIGKILL
Reason:       OOMKilled
Exit Code:    137

What this error actually means

The container runtime recorded an out-of-memory termination. Linux enforces container memory through control groups; when memory cannot be reclaimed within the applicable budget, the kernel can kill a process. Kubernetes records the terminated container state and, when its restart policy requires it, starts another instance. The new process has a clean memory baseline, which can hide the peak that killed its predecessor.

Exit code 137 alone is weaker evidence. It conventionally represents termination by signal 9, SIGKILL, but can also be an explicit exit status. A forced shutdown after the termination grace period, an operator’s kill, and an OOM kill can all end with 137. Read the termination reason, events and kernel evidence together before calling every such exit a memory leak.

A container limit and a node’s available memory are separate boundaries. A pod can hit its own limit while the host has gigabytes free, or become a victim of host memory pressure without reaching its configured limit. Increasing a request influences scheduling; it does not automatically raise an existing limit or repair allocations that grow without bound.

Causes, most common first

  1. 1Working set or concurrency exceeds the container budget. One request fits, but many concurrent requests retain buffers and query results simultaneously. A new batch size or traffic pattern changes peak resident memory without changing the steady-state graph. Rank this ahead of a leak when deaths align with specific jobs or bursts.
  2. 2Retained objects or queues grow with process age. Caches without eviction, queued retries and abandoned request objects accumulate. A rising baseline after idle periods distinguishes retention from temporary allocation. Increasing the limit postpones the same incident while increasing the eventual blast radius.
  3. 3Memory outside the managed heap was omitted. Native buffers, thread stacks, subprocesses and memory-backed emptyDir volumes can consume the same constrained budget. A heap dashboard is only one component of container memory; a flat managed heap cannot clear the rest of the process.
  4. 4The node ran out of usable memory. Overcommitted workloads or host processes can cause a system-level OOM event. Inspect the node before changing one workload’s limit. Kubelet eviction and kernel OOM killing are different paths and leave different evidence.

When you see it

  • The previous container state says OOMKilled while the current container is Running
  • Logs stop abruptly without the application’s normal shutdown sequence
  • Memory looks low immediately after each restart and climbs with accumulated work
  • A large import or several concurrent requests trigger kills between monitoring samples

How to diagnose it

Step 1

Preserve the previous termination record

Replace demo, app-pod and app with the affected namespace, pod and container. Read each container separately, including sidecars. Capture this before deleting the pod, because a replacement pod does not inherit its predecessor’s status.

kubectl -n demo get pod app-pod -o jsonpath='{range .status.containerStatuses[*]}{.name}{"\t"}{.lastState.terminated.reason}{"\t"}{.lastState.terminated.exitCode}{"\n"}{end}'
kubectl -n demo logs app-pod -c app --previous --timestamps

Step 2

Compare the admitted resource budget with events

Inspect the live object rather than only the deployment source: namespace defaults and admission policies may have changed its resources. Correlate the termination timestamp with probes, shutdown events and node conditions.

kubectl -n demo describe pod app-pod
kubectl -n demo get pod app-pod -o jsonpath='{.spec.containers[*].resources}{"\n"}'

Step 3

Read cgroup evidence where it is exposed

On a Linux cgroup v2 container with its cgroup mounted at this path, memory.events shows OOM counters and memory.max shows the enforced ceiling. A restarted container can have new counters; zero now does not disprove a previous kill. Use node-side retained telemetry for historical attribution.

kubectl -n demo exec app-pod -c app -- cat /sys/fs/cgroup/memory.events /sys/fs/cgroup/memory.current /sys/fs/cgroup/memory.max

Step 4

Capture peaks and classify the allocation shape

Graph container usage, concurrency and request sizes on the same timeline. kubectl top is a sampled snapshot, not a peak recorder. Reproduce the suspect workload with bounded concurrency and profile retained objects or native allocation before the kill occurs.

kubectl -n demo top pod app-pod --containers

The fix

Bound the multiplication that creates the peak: paginate results, stream uploads, cap worker concurrency and give queues finite capacity with an explicit overload response. Verify the total memory of simultaneous operations rather than measuring one successful request in isolation.

For retention, identify what still owns the objects after work completes and repair that lifecycle. For memory-backed files, bound and clean their contents. Moving a cache into a tmpfs volume moves the allocation mechanism without creating a new memory budget.

Raise a limit only when profiling establishes a legitimate working set and the node has capacity. Set a realistic request as well so the scheduler can place the workload honestly. Load-test the replacement under the actual limit; testing on an unrestricted workstation cannot validate the budget.

If 137 came from forced termination, repair signal delivery or draining instead. A larger memory limit cannot make an application observe SIGTERM, and it obscures the real correlation with deployment or scale-down events.

How to stop it coming back

  • Alert on restarts with termination reason and container name, not only pod availability
  • Record memory peaks and workload concurrency at a resolution that catches batch spikes
  • Exercise the maximum supported request size under the production memory limit
  • Keep memory requests, limits and runtime heap settings in the same capacity review

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Does exit 137 prove an OOM kill?

No. It commonly indicates SIGKILL, which has several causes. OOMKilled plus matching cgroup or kernel evidence establishes memory pressure much more strongly than the exit code alone.

Why is there no heap dump?

SIGKILL gives the runtime no opportunity to execute handlers or create a dump. Heap-dump-on-OutOfMemoryError only helps when the JVM itself throws that exception; the kernel can kill it before that happens.

Why did increasing the memory request not help?

A request reserves scheduling capacity and affects some pressure decisions. An explicitly configured container memory limit remains the same enforcement boundary until you change it.

Related

Other errors engineers hit next to this one

Full error and symptom index →