Java / JVM

java.lang.OutOfMemoryError: GC overhead limit exceeded

Written and reviewed by Sahil Srivastav

JVMGarbage collectionMemory
java.lang.OutOfMemoryError: GC overhead limit exceeded
	at java.base/java.util.HashMap.newNode(HashMap.java:1797)
	at java.base/java.util.HashMap.putVal(HashMap.java:630)

What this error actually means

This is the JVM refusing to die slowly. The collector noticed it had spent more than roughly 98% of recent wall-clock time in garbage collection while recovering less than 2% of the heap, concluded that progress had effectively stopped, and threw rather than letting the process thrash indefinitely.

It is therefore a *rate* diagnosis rather than a capacity one, and it is strictly more informative than plain "Java heap space". The heap is not simply full — it is full of objects that survive collection, so every collection cycle does maximum work for minimum reward. The application is alive, responding to health checks perhaps, and accomplishing nothing.

Two distinct situations produce it. Either the live set has grown to nearly fill the heap, so each collection can free almost nothing (a leak, or undersizing). Or the allocation rate is so high that the collector cannot keep pace even though nothing is leaking, which points at churn in a hot path rather than retention.

Causes, most common first

  1. 1Live set nearly equal to heap size. Whether from a leak or from genuine data growth, when survivors fill most of the heap every collection traces almost everything and frees almost nothing. This is the most common shape and is diagnosed from post-full-GC occupancy.
  2. 2Extreme allocation rate in a hot path. Boxing in a tight loop, string concatenation per row, a new buffer or temporary collection per element, defensive copies inside an inner loop. Nothing is retained, so a heap dump looks clean — but the collector is pinned servicing churn.
  3. 3Premature promotion from an undersized young generation. Short-lived objects survive young collections only because the survivor space cannot hold them, so they get promoted into the old generation and must be collected there — the most expensive possible way to handle garbage that should have died young.
  4. 4Humongous allocations under G1. Objects larger than half a G1 region are allocated directly into the old generation across contiguous regions. A stream of large arrays or buffers fragments the heap and forces full collections regardless of total occupancy.
  5. 5Heap sized below the real working set. The simple case, and the only one where raising `-Xmx` is the right answer. Identified by a stable, low post-collection live set that is nonetheless a large fraction of a small heap.

When you see it

  • Throughput collapses to near zero while CPU sits at 100%, almost entirely in GC threads
  • Latency percentiles blow out minutes before the error, with p99 in the tens of seconds
  • GC log shows back-to-back full collections that reclaim a few megabytes each
  • Health checks may keep passing, so the orchestrator does not replace the instance
  • It often follows a deploy that added an allocation in a per-row or per-request loop

How to diagnose it

Step 1

Measure the GC time fraction

Confirm the diagnosis in numbers before changing anything: total GC time as a proportion of uptime, and bytes reclaimed per full collection. Under 5% is healthy; above 30% the application is effectively down.

jstat -gcutil <pid> 1s 20

Step 2

Read post-full-GC occupancy over time

This single series separates leak from churn. Rising across hours means retention. Flat and low while GC time is high means allocation rate — the collector is busy but the survivors are few.

-Xlog:gc*:file=/var/log/app/gc.log:time,uptime,level,tags:filecount=5,filesize=20M

Step 3

Find the allocation hot path if survivors are few

Async-profiler in allocation mode attributes bytes to stack traces with low overhead, which is what you need when the problem is churn and a heap dump looks innocent.

asprof -e alloc -d 30 -f alloc.html <pid>

Step 4

Check promotion and region behaviour

Look for survivor-space overflow and "to-space exhausted" or humongous allocation lines in the GC log. These name the mechanism directly and point at sizing rather than at your object graph.

grep -E "Humongous|to-space|Pause Full" /var/log/app/gc.log | tail -40

The fix

If survivors are growing, this is the heap-space problem in a different costume: find the retention with a heap dump, bound the cache or collection, and fix the lifecycle. Do not tune the collector to accommodate a leak.

If survivors are few and allocation rate is high, cut allocations in the hot path specifically. Hoist buffers out of loops and reuse them, replace boxed collections with primitive-specialised ones, avoid string concatenation per row in favour of a single builder, and stop making defensive copies inside inner loops. The target is bytes allocated per operation, and it is usually improvable by an order of magnitude without changing behaviour.

If the GC log shows survivor overflow, give the young generation and survivor spaces room to let short-lived objects die young: raise the young-gen size and check `MaxTenuringThreshold` rather than raising total heap. Premature promotion is a sizing bug, not a code bug.

For humongous allocations under G1, either shrink the objects below half a region or raise `-XX:G1HeapRegionSize` so they are no longer humongous. Streaming large payloads in chunks removes the problem entirely.

Only with a stable, low live set and low allocation rate should you raise `-Xmx`. And never disable the check with `-XX:-UseGCOverheadLimit` — that converts a clear error into an indefinite hang, which is strictly harder to diagnose and worse for users.

How to stop it coming back

  • Track allocation rate (bytes/sec) as a first-class metric alongside heap used; regressions show up here one deploy before they show up as incidents
  • Alarm on GC time fraction above 10% sustained, which gives minutes of warning before collapse
  • Include a GC-log review in performance-sensitive pull requests that touch per-row or per-request loops
  • Fail health checks when GC time fraction is pathological so the orchestrator can replace a thrashing instance
  • Benchmark hot paths with allocation profiling in CI rather than wall-clock timing alone

Practise this failure in a real repository

Gronex ships a service whose hot path allocates enough to pin the collector while leaking nothing at all — the case a heap dump cannot solve. You get GC evidence and an allocation profile, and the tests assert bytes allocated per operation, not wall-clock time.

FAQ

How is this different from "Java heap space"?

Same underlying pressure, earlier and more informative detection. "Java heap space" means one allocation failed. This means the collector measured its own futility — over 98% of time spent for under 2% reclaimed — and gave up. If you see it, you have a rate problem to characterise, not just a full heap.

Is disabling UseGCOverheadLimit ever right?

Almost never in a service. It removes the error without removing the thrashing, so instead of a crash you get an instance that is permanently unavailable but reports itself alive. Batch jobs willing to trade hours of wall clock for completion are the only defensible case.

Can this happen with no memory leak at all?

Yes, and it is a frequently missed diagnosis. A high enough allocation rate saturates the collector with entirely short-lived objects. The heap dump looks clean, post-GC occupancy is flat, and the fix is in the hot path rather than in any object’s lifetime.

Will switching collectors fix it?

It can change how the failure presents — ZGC or Shenandoah trade throughput for pause time — but a collector cannot create memory that does not exist or keep up with an unbounded allocation rate. Characterise the cause first; collector choice is a refinement, not a fix.

Related

Other errors engineers hit next to this one

Full error and symptom index →