Linux / shell
Out of memory: Killed process — find which memory boundary was exhausted
Written and reviewed by Sahil Srivastav
Out of memory: Killed processWhat this error actually means
The displayed text is the stable prefix of a kernel OOM victim message; a real log continues with a PID, process name and memory figures. It means the kernel selected a process for termination while handling memory exhaustion. The killed application may leave no exception or shutdown log because SIGKILL cannot be caught to run cleanup.
There are two scopes to separate immediately. A host can run out of reclaimable memory, or a memory cgroup can hit its own enforced boundary while the host still has free capacity. Host-level free output alone cannot rule out a container or service limit. Conversely, a low application heap limit is not proof that its whole process fits inside its cgroup.
Exit status 137 is commonly a shell representation of termination by signal 9, but it does not identify who sent that signal. An operator, supervisor or timeout can also send SIGKILL. Require a matching kernel event, runtime status or cgroup evidence before calling an unexplained kill an OOM incident.
Causes, most common first
- 1An operation materialises input without a bound. A report, join or query result can consume memory proportional to tenant data rather than normal request size. Several concurrent requests multiply that peak. Small development fixtures therefore do not establish a safe production bound.
- 2The limit covers more than the application heap. Native allocations, thread stacks, buffers, child processes and charged file cache can consume cgroup memory alongside the managed heap. Setting a heap ceiling equal to the container limit leaves no room for these other consumers.
- 3Retained state grows across completed requests. Unbounded caches, queues and long-lived references push the baseline upwards. Raising the limit may lengthen the time to failure while allowing a larger backlog. Observe the baseline after traffic drains to distinguish retention from a transient peak.
- 4Host or ancestor-cgroup pressure constrains the workload. A sibling process may consume shared capacity, and an ancestor cgroup can impose a tighter collective budget. Inspect the hierarchy and other workload demand instead of attributing every host kill to the largest heap graph you can see.
When you see it
- The process disappears without a language-level stack trace or graceful shutdown
- A container repeatedly dies during imports, report generation or startup warmup
- The host has available memory but a limited workload still gets killed
- Heap metrics look safe while total process or cgroup memory keeps growing
How to diagnose it
Step 1
Correlate kernel evidence with the incident time
On a systemd host, kernel journal access may require elevated read permissions. Search around the kill timestamp and inspect surrounding lines for the constrained scope and victim. A container often cannot read host kernel logs, so collect them from the node.
journalctl -k --since '30 minutes ago' | rg -i 'out of memory|oom|killed process'Step 2
Locate the process’s memory cgroup
Use a still-running instance or capture this before reproducing. On a unified v2 hierarchy, the 0:: entry identifies the path relative to the cgroup mount as seen in that namespace. Do not blindly append a host path inside a differently rooted container.
app_pid=1234
cat /proc/"$app_pid"/cgroup
findmnt -t cgroup2Step 3
Read counters and limits at the relevant cgroup
Replace this example path with the resolved v2 cgroup directory. Compare event counters before and after an incident: historical nonzero values alone do not date the failure. Inspect constrained ancestors too; memory.events can include descendant activity.
app_cgroup=/sys/fs/cgroup/system.slice/app.service
cat "$app_cgroup/memory.current" "$app_cgroup/memory.max"
cat "$app_cgroup/memory.events"
cat "$app_cgroup/memory.stat"Step 4
Compare total memory with application metrics
Capture these while the workload is alive, then align with heap and request-concurrency measurements. RSS is useful but does not equal all cgroup charges. Break down anonymous and file-backed usage in memory.stat before blaming managed objects alone.
ps -p 1234 -o pid,rss,vsz,comm
free -hThe fix
Bound the work that creates the peak: stream rows, limit export ranges, cap in-flight jobs and avoid holding a full input and output simultaneously. Test with the largest supported input at the supported concurrency, since a one-request benchmark understates the required budget.
For retention, find the owner keeping data alive and introduce explicit bounds or release points. Capture language-level profiles before memory is exhausted. A kernel kill may occur before an application can generate its usual out-of-memory dump, so do not rely solely on an exception-triggered diagnostic.
Size the whole workload after the leak or unbounded operation is addressed. Leave measured headroom beyond the heap for native memory and other cgroup consumers. Coordinate the service limit with host capacity and sibling workloads; raising one limit is not a substitute for available physical capacity.
Recover through the supervisor after preserving evidence and reducing the triggering load. Make interrupted jobs safe to resume: SIGKILL can land between an external side effect and recording its completion. Memory recovery should not create duplicate business operations on restart.
How to stop it coming back
- Alert on memory pressure and remaining cgroup headroom before a kill; retain event counters and kernel logs across restarts.
- Exercise maximum input size and concurrency together, including startup and background jobs.
- Track process and cgroup memory alongside managed-heap metrics so native growth is visible.
FAQ
Does exit 137 always mean OOM?
No. It commonly indicates SIGKILL, which has several possible senders. Correlate the time and PID with kernel or runtime evidence rather than inferring the cause from the exit code alone.
Why was a small process killed?
Victim selection is not simply a sort by the application metric you are viewing. The constrained scope and OOM selection policy matter. Read the surrounding kernel report to identify the exhausted boundary before tuning a different workload.
Can I handle SIGKILL to flush data?
No. Design persistence and retries to survive abrupt termination. Graceful shutdown handlers are useful for planned stops but cannot provide durability when the kernel kills the process.
Related
Other errors engineers hit next to this one
- No space left on device despite free disk space
- Text file busy during executable replacement
- set -e script continues after a failed pipeline
- An unquoted variable turns one argument into several
- OOMKilled — container exit code 137
- CrashLoopBackOff
- ImagePullBackOff / ErrImagePull
- Readiness probe failed