Java / JVM

java.lang.OutOfMemoryError: Metaspace

Written and reviewed by Sahil Srivastav

JVMClassloadersMemory
java.lang.OutOfMemoryError: Metaspace
	at java.base/java.lang.ClassLoader.defineClass1(Native Method)
	at java.base/java.lang.ClassLoader.defineClass(ClassLoader.java:1017)

What this error actually means

Metaspace holds class metadata — the runtime representation of every class the JVM has loaded — in native memory, outside the heap. It fills when classes are loaded faster than they are unloaded. And a class can only be unloaded when its entire defining classloader becomes unreachable.

That last sentence is the whole problem. Classloaders are unloaded as a unit, so a single lingering reference to one object defined by a classloader pins every class that loader ever defined, along with their static fields and everything those reference. One stray reference retains megabytes of metadata permanently.

This is why Metaspace exhaustion is almost never about sizing. Unlike heap, the working set of class metadata in a stable application is essentially constant — it loads at startup and stops. Growth over time is a leak by definition, and the growth is nearly always classloaders that should have died.

Causes, most common first

  1. 1A classloader pinned by a stray reference after reload. The canonical cause. A plugin or web application is undeployed, but something outside it still holds one object it defined — a `ThreadLocal` on a pooled thread, a JDBC driver registered in `DriverManager`, a listener in a static registry, a shutdown hook, a running timer thread. The loader never dies, so none of its classes unload.
  2. 2Runtime proxy or bytecode generation in a loop. Dynamic proxies, CGLIB subclasses, Groovy or scripting-engine compilation, or expression evaluation generating a fresh class per invocation instead of per type. Each generated class is real metadata that accumulates forever if the generator caches nothing.
  3. 3A thread started by the reloaded code and never stopped. A live thread whose `Runnable` or context classloader belongs to the old deployment is a GC root pointing straight at that classloader. This also silently keeps old application behaviour running alongside the new one.
  4. 4Deserialisation or reflection caches keyed unboundedly. Frameworks that cache per-type accessors, serialisers, or generated adapters can accumulate classes when the type space is effectively unbounded — for example a class generated per dynamic schema version.
  5. 5A genuinely large class space with a low explicit limit. Real but rare and easy to identify: class count plateaus and stays flat, but the plateau is above a `MaxMetaspaceSize` that was set conservatively. Typical for very large applications or heavy framework stacks.

When you see it

  • Loaded class count grows without bound while the application does the same work
  • It appears only after the Nth hot redeploy, plugin reload, or module refresh — never on a cold start
  • Raising `-XX:MaxMetaspaceSize` moves the failure later by a predictable multiple and changes nothing else
  • Full GCs become frequent and long as the JVM tries repeatedly to unload classes it cannot
  • Container RSS grows while heap usage stays flat, since Metaspace is native memory

How to diagnose it

Step 1

Confirm it is growth, not size

Count loaded classes over time. Continuous growth while the workload is constant proves a leak; a plateau proves undersizing. This one measurement decides everything that follows.

jcmd <pid> VM.classloader_stats
jstat -class <pid> 5s

Step 2

Find the classloaders that should be dead

List loaders and look for several instances of the same application or plugin loader. N copies means N reloads have leaked. The loader hierarchy dump shows exactly how many generations are still alive.

jcmd <pid> GC.class_stats
jcmd <pid> VM.classloaders show-classes

Step 3

Trace the path from a GC root to the dead loader

Take a heap dump and, in MAT, find the classloader instance then run "Path to GC Roots" excluding weak and soft references. That path is the bug, and it is usually one line long: a static map, a `ThreadLocal`, or a thread.

jcmd <pid> GC.heap_dump /tmp/meta.hprof

Step 4

Rule in proxy generation

If class names in the histogram contain `$Proxy`, `$EnhancerBy`, `GeneratedConstructorAccessor`, or a script-engine prefix, and the count rises with request volume, you are generating classes per call rather than per type.

jmap -histo:live <pid> | grep -E "Proxy|Enhancer|Lambda" | head

The fix

For the pinned classloader, break the single reference the dump identified. In practice this means: clear every `ThreadLocal` set by the reloadable code in a `finally` block, deregister JDBC drivers and MBeans on shutdown, remove listeners from static registries symmetrically, and stop and join every thread the module started before the module is discarded.

For proxy or bytecode generation, cache the generated class per type rather than per invocation, and make the cache key the type — not the instance and not the request. If a scripting engine compiles user input, compile once per distinct script and cache by content hash with a bound.

For threads, never let reloadable code start an unmanaged thread. Use an executor owned by the container so its lifecycle is tied to deployment, and shut it down with a timeout during undeploy.

If and only if class count plateaus, raise the limit deliberately: `-XX:MaxMetaspaceSize=512m` and keep it set rather than unbounded, so a future leak fails fast and loudly rather than consuming the container’s memory and getting the pod OOM-killed with no Java-level error at all.

How to stop it coming back

  • Alarm on loaded class count trend, not Metaspace bytes — growth is the signal
  • Treat hot redeploy as a tested code path: redeploy 20 times in CI and assert class count returns to baseline
  • Always set `MaxMetaspaceSize` explicitly. Unbounded Metaspace converts a diagnosable Java error into an opaque container kill
  • Enforce symmetric register/deregister in review for anything touching a static or thread-scoped registry
  • Prefer restarting a process over hot reloading in production — the vast majority of these leaks simply cannot occur

Practise this failure in a real repository

Gronex ships a plugin host that leaks a classloader on every reload. You get the class-count evidence and the loader dump, and the test suite reloads repeatedly and asserts the old loader becomes unreachable — so only genuinely breaking the retention path passes.

FAQ

Is Metaspace part of the heap?

No. It is native memory outside the heap, which is why heap graphs look perfectly healthy while this error fires, and why an unbounded Metaspace can get your container OOM-killed by the kernel without any Java error being logged at all.

Why did removing PermGen not solve this?

Java 8 replaced PermGen with Metaspace, which grows into native memory by default rather than having a fixed small limit. That changed the default failure mode from a quick error to slow native-memory growth. The underlying classloader-leak mechanism is unchanged.

How do I know which reference is pinning the loader?

"Path to GC Roots" on the classloader instance in a heap dump, with weak and soft references excluded. The remaining path is the leak. Suspect, in order: `ThreadLocal` on a pooled thread, a live thread, a static collection, a JDBC driver, an MBean.

Do lambdas leak Metaspace?

Each lambda call site generates a class, but the count is bounded by the number of call sites in your code, so it is a fixed startup cost. Unbounded lambda class growth means you are generating call sites at runtime — that is a different bug, not lambdas themselves.

Related

Other errors engineers hit next to this one

Full error and symptom index →