Linux / shell
JVM memory exceeds the container limit despite a healthy heap
Written and reviewed by Sahil Srivastav
Reason: OOMKilled
Exit Code: 137What this error actually means
The Java heap is one allocation arena inside a larger process. A container memory limit covers more than live Java objects: committed heap pages, metaspace, code cache, thread stacks, direct buffers and other charged memory can all contribute. Setting Xmx equal to the container limit leaves no room for the rest of the process.
MaxRAMPercentage controls heap sizing ergonomics relative to memory recognised by the JVM; it is not a cap on total resident memory. An explicit Xmx already determines the maximum heap, so adding a percentage flag does not shrink that explicit maximum. Read the effective flags in the running process instead of assuming an environment variable changed the selected heap.
When the kernel kills the process first, Java cannot throw OutOfMemoryError or run a shutdown hook. This explains an apparently comfortable heap graph followed by OOMKilled and no heap dump. Container awareness helps the JVM choose a heap, but it cannot infer how many native buffers, threads or subprocesses the application will need.
Causes, most common first
- 1Heap maximum consumes almost the whole container budget. A deployment sets Xmx to the same value as the limit or uses a very high percentage copied from another workload. Even a stable heap then competes with mandatory runtime structures. There is no universal safe percentage independent of thread count, libraries and workload.
- 2Unbounded native or direct-buffer demand. Networking libraries, compression, JNI and memory mappings can grow outside ordinary object occupancy graphs. Heap wrappers may be small while their native backing allocations are large. The memory retained by one connection multiplied by connection concurrency is often the useful capacity model.
- 3Thread count or class metadata grows. Platform threads require native stack resources, and repeated class loading consumes metadata. Raising a heap maximum does not reduce either cost. Inspect growth over time to distinguish legitimate worker capacity from leaked executors or class loaders.
- 4The effective runtime configuration differs from the intended one. An entrypoint adds Xmx, a base image supplies JAVA_TOOL_OPTIONS, or a runtime build does not interpret the host’s cgroup setup as expected. Configuration visible in a manifest is only input; the selected MaxHeapSize and detected limit are the result to verify.
When you see it
- Container memory approaches its limit while heap occupancy remains below Xmx
- The process disappears without OutOfMemoryError or HeapDumpOnOutOfMemoryError output
- Increasing Xmx makes container kills happen sooner rather than improving stability
- More concurrent connections or threads increase memory without a comparable live-heap increase
How to diagnose it
Step 1
Confirm memory termination and the actual container limit
Use the affected namespace, pod and container. Confirm OOMKilled rather than assuming every 137 is memory-related. Record the image digest and Java version, because container detection depends on the runtime build and operating environment.
kubectl -n demo describe pod app-pod
kubectl -n demo exec app-pod -c app -- java -versionStep 2
Inspect the running JVM’s selected heap
The commands assume Java is PID 1 and jcmd is installed; substitute its actual PID and use the same effective user. VM.flags reveals the selected MaxHeapSize and relevant flags. A new java -version process does not tell you the running application’s effective heap.
kubectl -n demo exec app-pod -c app -- jcmd 1 VM.flags
kubectl -n demo exec app-pod -c app -- jcmd 1 GC.heap_infoStep 3
Compare cgroup accounting with native categories
On cgroup v2, memory.current is the group’s charged usage. Native Memory Tracking must be enabled when starting the JVM with -XX:NativeMemoryTracking=summary. Its committed values help attribute JVM memory, but do not equal RSS and do not account for every third-party native allocation.
kubectl -n demo exec app-pod -c app -- cat /sys/fs/cgroup/memory.current /sys/fs/cgroup/memory.max
kubectl -n demo exec app-pod -c app -- jcmd 1 VM.native_memory summary scale=MBStep 4
Compare growth across a representative interval
With NMT enabled, set a baseline before the workload, then request a difference after it. Match native category growth to thread counts, buffer metrics and container usage. Flat NMT with rising container memory points towards untracked native allocations, charged file memory or other processes rather than proving no leak exists.
kubectl -n demo exec app-pod -c app -- jcmd 1 VM.native_memory baseline
# Run the representative workload between these commands.
kubectl -n demo exec app-pod -c app -- jcmd 1 VM.native_memory summary.diff scale=MBThe fix
Construct a complete memory budget from measured peaks: managed heap, non-heap JVM memory, library buffers, other container processes and safety headroom. Choose Xmx or a percentage policy that leaves that headroom under the enforced limit. Test at peak concurrency and during warm-up, not only after the service becomes idle.
Remove conflicting heap settings from launch scripts and environment injection. If you use MaxRAMPercentage, confirm that the runtime detects the correct constrained memory and that an explicit Xmx does not override your intent. Modern supported JDKs provide container logging such as -Xlog:os+container=info for a controlled diagnostic start.
Bound native demand at its owner: cap connection concurrency and retained buffers, close resources, and shut down abandoned executors. A direct-memory limit can turn one allocation class into a visible failure, but it does not constrain all JNI allocations or replace an application overload policy.
For a legitimate working set that cannot fit, increase the container budget together with realistic scheduling requests and node capacity. Do not merely lower the heap until the kernel stops killing Java if that forces continuous GC and makes the service unusable. Validate latency, throughput and memory together.
How to stop it coming back
- Export heap, non-heap, thread count and container memory on the same dashboard
- Review JVM flags and container limits as one configuration change
- Capture native-memory baselines in staging before long-running load tests
- Repeat memory tests when native libraries, JDK builds or thread models change
FAQ
Is MaxRAMPercentage=75 always safe?
No. The remaining quarter may be ample for one process and inadequate for another. The non-heap requirement depends on absolute memory, concurrency, loaded classes and native libraries. Measure it under the enforced limit.
Why does NMT not match the container graph?
NMT reports tracked JVM allocation categories and reserved or committed memory. Container accounting measures charged memory across a broader boundary. Third-party native code, file-backed memory and other processes can explain part of the difference.
Will HeapDumpOnOutOfMemoryError catch a container kill?
Not when SIGKILL occurs first. That option reacts to a Java OutOfMemoryError. Capture profiles and metrics before the cgroup reaches its limit, and distinguish a JVM exception from an operating-system kill.
Related
Other errors engineers hit next to this one
- Thread pool starvation — every worker waiting on a task in its own pool
- Partially constructed object published by double-checked locking
- Lost update from get-then-put on a ConcurrentHashMap
- CompletableFuture failed with nothing logged
- awaitTermination never returns and the JVM will not exit
- InterruptedException caught and ignored — the task can no longer be cancelled
- Two unrelated components sharing a monitor via a boxed Integer or interned String
- Cache stampede — the same expensive value built many times concurrently