Java / JVM

java.lang.OutOfMemoryError: Java heap space

Written and reviewed by Sahil Srivastav

JVMMemoryProduction incident
Exception in thread "http-nio-8080-exec-7" java.lang.OutOfMemoryError: Java heap space
	at java.base/java.util.Arrays.copyOf(Arrays.java:3537)
	at java.base/java.lang.StringBuilder.append(StringBuilder.java:242)

What this error actually means

The JVM attempted to allocate an object, found insufficient contiguous space in the heap, and could not reclaim enough by collecting. The critical detail is that the stack trace names the *victim* — the allocation that happened to be unlucky — not the cause. The thread printed in the trace is very often innocent.

There are three distinct failure shapes behind this one message, and they are separated by how heap usage behaves over time rather than by the trace. A leak shows live-set growth that survives every full collection, sawtoothing upward until the ceiling. An unbounded single operation shows a normal baseline with one vertical spike. Genuine undersizing shows a stable, healthy sawtooth that simply sits too close to the limit.

Only the third shape is fixed by `-Xmx`. For the first two, raising the heap changes nothing except how long you wait for the outage — and it makes the eventual pause longer and the heap dump harder to analyse.

Causes, most common first

  1. 1Unbounded collection used as a cache. A `HashMap` or `ConcurrentHashMap` keyed by something unbounded — session id, tenant, request key — with entries added and never evicted. It works perfectly in testing because tests do not run for four days.
  2. 2A query that materialises the whole result set. An endpoint fetches every row for an entity rather than a page, then filters in memory. Fine for the developer’s seed data, fatal for the customer with 400,000 orders. This is the classic "works in dev, dies for one tenant" shape.
  3. 3Registered listeners or callbacks never removed. Anything that registers into a long-lived registry — event listeners, metrics gauges, shutdown hooks, `ThreadLocal` values on pooled threads — keeps the registering object, and its whole reference graph, reachable forever.
  4. 4Streaming response buffered fully before sending. Building a report or export into a `StringBuilder`, a byte array, or a list before writing to the response. Heap cost scales with output size and with concurrency, so two simultaneous exports can double what one instance survived.
  5. 5Genuinely insufficient heap for the working set. Real but much rarer than assumed, and only diagnosable after ruling the others out. Typical after a legitimate data-volume increase or moving to a container with a lower memory limit than the JVM believes it has.

When you see it

  • The service degrades for minutes before dying, with GC CPU climbing as throughput falls
  • Restarting buys a predictable amount of time — hours, or a fixed number of requests
  • The reported thread differs every time, because the victim is random
  • One specific endpoint or a specific tenant’s data correlates with the crash
  • Old-generation occupancy after full GC rises monotonically across days

How to diagnose it

Step 1

Always capture the dump automatically

Configure this before you need it. Without a dump from the moment of failure, you are guessing — and the failure will not reproduce on demand.

-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/log/app/heap.hprof

Step 2

Find what is big, then find who holds it

Open the dump in Eclipse MAT and run the Leak Suspects report. The useful answer is never the class with the most instances — it is the shortest path from a GC root to the dominator tree of the biggest retained set.

jmap -histo:live <pid> | head -25

Step 3

Classify the shape from the GC log, not the trace

Look at old-gen occupancy immediately *after* each full collection. Rising monotonically means a leak. Flat with one spike means a single unbounded operation. Flat and high means undersizing.

-Xlog:gc*,gc+heap=debug:file=/var/log/app/gc.log:time,uptime:filecount=5,filesize=20M

Step 4

Correlate with input size

If the dump points at a result set or export buffer, query production for the largest instance of that input — the biggest account, longest history, widest date range. A leak is independent of input size; an unbounded operation is proportional to it.

The fix

For an unbounded cache, give it an explicit bound and eviction policy. `LinkedHashMap` in access-order mode with a `removeEldestEntry` override is enough for a local cache; anything with real requirements should use Caffeine with a size or weight limit. A cache without a bound is not a cache, it is a leak with good intentions.

For a materialising query, push the limit down to the database. Paginate with keyset pagination rather than `OFFSET`, and stream rather than collect when the result is genuinely large. The invariant to enforce is that memory per request is bounded by page size, not by the customer’s total history.

For registration leaks, pair every register with a deregister and make the deregister run on the failure path too. On pooled threads, always clear `ThreadLocal` values in a `finally` block — the thread outlives the request, so whatever you leave there is retained until the pool shuts down.

For buffered responses, write incrementally to the output stream and flush periodically, keeping only the current chunk in memory.

Only if the dump shows a healthy, bounded live set should you raise `-Xmx`. In a container, also set `-XX:MaxRAMPercentage` so the JVM sizes itself against the cgroup limit rather than the host’s total memory.

How to stop it coming back

  • Alarm on old-gen occupancy after full GC, not on heap used — the former is the leak signal and the latter is noise
  • Make every cache declare a maximum size at construction; treat an unbounded map in a long-lived object as a review blocker
  • Add tests that assert bounded memory behaviour: N pages returns O(page) objects regardless of total rows
  • Ship `-XX:+HeapDumpOnOutOfMemoryError` in every environment from day one
  • Load-test with production-shaped data volumes, especially the largest single tenant, rather than uniform seed data

Practise this failure in a real repository

Gronex gives you the heap evidence and a repository that leaks: a cache with no bound, a retention path through a long-lived registry, and a test suite that asserts the live set stays bounded under sustained load. Raising the heap does not make it pass.

FAQ

Why does the stack trace point at code that looks harmless?

Because the trace shows whichever allocation happened when the heap ran out, not the allocation that filled it. A `StringBuilder.append` in the trace almost never means the string building is the bug. Use the heap dump for cause and the trace only for timing.

Can the JVM throw this and keep running?

Sometimes, and that is worse than crashing. A thread can die mid-operation leaving a lock held, a half-written record, or a queue permanently drained of its consumer. Prefer `-XX:+ExitOnOutOfMemoryError` so the orchestrator replaces a process whose state you can no longer trust.

Does this mean I have a memory leak?

Not necessarily. One unbounded request against unusually large data produces the identical message with no leak present. The GC log tells you which: post-full-GC occupancy rising over days is a leak, a single vertical spike from a flat baseline is not.

Why did it start only after moving to Kubernetes?

On older JVMs, or without `MaxRAMPercentage`, the JVM can size the heap against the node’s memory instead of the container limit, so the kernel OOM-kills the pod or the heap targets a ceiling it can never reach. Set the percentage explicitly and verify with `java -XX:+PrintFlagsFinal -version | grep MaxHeapSize`.

Related

Other errors engineers hit next to this one

Full error and symptom index →