Concurrency
A ThreadLocal value retained on a pooled thread
Written and reviewed by Sahil Srivastav
jmap -histo:live 4412 | head -12
num #instances #bytes class name (module)
-------------------------------------------------------
1: 200 838860800 [B (java.base@21)
2: 200 9600 com.example.audit.RequestContext
3: 200 6400 java.lang.ThreadLocal$ThreadLocalMap$Entry
# and the retention path from the heap dump:
java.lang.Thread @ 0x70a1c2f40 "http-nio-8080-exec-31"
|- threadLocals java.lang.ThreadLocal$ThreadLocalMap @ 0x70a1c3118
|- table java.lang.ThreadLocal$ThreadLocalMap$Entry[16] @ 0x70a1c3180
|- [7] ThreadLocalMap$Entry @ 0x70a1c31e8
|- value com.example.audit.RequestContext @ 0x70a1c3220 (retains 4.1 MB)What this error actually means
A `ThreadLocal` is not scoped to your request, your transaction, or your call stack. It is scoped to the `Thread` object, and a pooled thread outlives every request it serves. Whatever you leave in it is still there when the pool hands that thread to the next caller.
The storage is a field on `Thread` itself — `threadLocals`, a `ThreadLocalMap` whose entries use a **weak** reference to the `ThreadLocal` key and a **strong** reference to the value. People read "weak" and assume the whole entry is collectable. It is not: when the key is collected the entry becomes stale, but the value is still strongly reachable from the live thread until some later `get`, `set`, or `remove` on that same thread happens to walk past the stale slot and clear it. On a thread that goes idle in a pool, that never happens.
So one missing `remove()` produces two different bugs from the same cause, and which one you notice first is luck. The **correctness** bug is data bleed: request B on the same worker sees request A’s tenant id, user principal, locale, or trace id, because the value was never cleared and B’s code path did not set it. The **memory** bug is retention: the value, and its entire reference graph, is pinned for the lifetime of the pool — and if that graph reaches a classloader, you get Metaspace growth instead of heap growth.
The severity asymmetry is worth naming. A leaked megabyte is an operational problem. A leaked security principal is a cross-tenant data exposure that your tests, which run one request at a time on a fresh thread, cannot see at all.
Causes, most common first
- 1A filter or interceptor that sets but never removes. The overwhelmingly common shape. Context is populated at the start of a request and the removal is either absent or placed after the handler call rather than in a `finally`, so any exception or early return leaves it set. The next request on that thread inherits it.
- 2Removal on the success path only. The `remove()` exists but sits at the end of the `try` body. Every thrown exception skips it. Because exceptions are rare, the leak accumulates slowly — one thread poisoned per failure — which is why it looks intermittent rather than systematic.
- 3Context set on a thread that did not start the request. Work is handed to an executor, a `CompletableFuture`, or a parallel stream, and the context is copied onto the downstream thread. Now two threads hold it and only one cleanup runs. Parallel streams are the nastiest case because they borrow the common ForkJoinPool, whose threads live for the whole JVM.
- 4A value that transitively reaches something enormous. The retained object graph, not the value, is what matters. A context holding a fetched entity, an open buffer, a `ResultSet` row list, or a Spring bean can retain megabytes. A context holding a class defined by a reloadable classloader pins that loader and every class it ever defined.
- 5Library `ThreadLocal`s on threads the library does not own. Formatters, JDBC diagnostic contexts, serialiser caches, and `SimpleDateFormat` holders are frequently cached per thread deliberately. That is fine on a thread whose lifetime the library controls, and a leak on an application pool the library merely borrowed — especially across a redeploy.
- 6`InheritableThreadLocal` copied into pool threads at creation. The value is snapshotted when the child thread is created. A pool thread created during request 1 keeps request 1’s inherited value for its entire life, serving it to every later request, and no cleanup in request 1 can undo that.
When you see it
- A request occasionally sees another user’s tenant, locale, or feature flags — and always on a warm instance, never on a cold one
- The wrong trace id or MDC field appears in logs for a subset of requests, corrupting correlation
- Heap grows in proportion to pool size rather than to traffic, then plateaus at roughly `poolSize × valueSize`
- Old-generation occupancy after full GC includes `ThreadLocal$ThreadLocalMap$Entry` instances that never fall
- A heap dump shows exactly as many copies of the context object as there are live pool threads
- The bug rate correlates with how *unevenly* work is distributed across threads, so it appears and vanishes as traffic shifts
How to diagnose it
Step 1
Assert the invariant at the pool boundary
The decisive test is not a memory measurement, it is an assertion at the point where a thread is handed back. Wrap the executor so that after every task it verifies the holder is empty, and fail loudly. This converts an intermittent production symptom into a deterministic CI failure.
if (RequestContext.current() != null) {
throw new IllegalStateException("context leaked on " + Thread.currentThread().getName());
}Step 2
Count the retained copies in a live heap
One instance per live pool thread, unchanging while traffic flows, is the signature. Traffic-proportional growth is a different bug — that is an unbounded collection somewhere else.
jcmd <pid> GC.class_histogram | grep -E "RequestContext|ThreadLocalMap"Step 3
Walk the retention path in a dump
Take a dump and, in Eclipse MAT, select the value type and run "Path to GC Roots" with weak references excluded. A path of `Thread → threadLocals → table[i] → value` is conclusive and also tells you which pool the thread belongs to from its name.
jcmd <pid> GC.heap_dump /tmp/tl.hprofStep 4
Reproduce the data bleed deterministically
Pin the pool to a single thread and run two requests in sequence where the second does not set the context. If the second observes the first’s value, the bleed is proven. A one-thread executor makes a race into a certainty.
var pool = Executors.newFixedThreadPool(1); // forces thread reuseStep 5
Look for the value on threads you did not expect
Grep the dump for the value type among ForkJoinPool worker threads. Hits there mean a parallel stream or an async stage propagated the context onto a JVM-lifetime thread, where no request-scoped cleanup will ever run.
grep -c "ForkJoinPool.commonPool-worker" dump.txtThe fix
Make `remove()` unconditional. Every `set()` needs a matching `remove()` in a `finally` on the same thread, in the same method, at the same nesting level. `remove()` — not `set(null)`: setting null leaves a live entry holding a null value, so the map slot survives and the `get()` returns null in a way that is indistinguishable from "never set", which hides the next bug.
Own the boundary rather than trusting every caller. Expose the context only through a scoped helper that acquires and releases around a lambda, so no calling code can forget the cleanup. This is the single highest-value change, because it removes the failure mode instead of documenting it.
For work handed to another thread, propagate a copy explicitly and clean it up on that thread too. A wrapping `Executor` decorator that captures on submit, installs before `run`, and removes in `finally` is the only shape that survives exceptions, cancellations, and rejected tasks.
Never let a context object hold a large or long-lived graph. Put identifiers in it, not entities: a tenant id rather than a tenant, a user id rather than a loaded principal with its permission tree. Then even a genuine leak costs bytes rather than megabytes.
On Java 21 and later, prefer `ScopedValue` for request-scoped data where you can. It is bound for the duration of a dynamic scope and unbound on exit by construction, so there is no cleanup to forget and no way for a value to outlive its scope onto a pooled thread.
Avoid `InheritableThreadLocal` entirely with pools. Inheritance happens at thread creation, which for a pool is an arbitrary moment unrelated to your request boundaries, so the semantics you want are not available.
// Leaks: remove() runs only when the handler returns normally
CTX.set(RequestContext.from(request));
chain.doFilter(request, response); // throws -> value stays on the pool thread
CTX.remove();
// Correct: cleanup is unconditional
CTX.set(RequestContext.from(request));
try {
chain.doFilter(request, response);
} finally {
CTX.remove(); // not set(null): that keeps a live entry
}
// Better: the boundary owns the lifecycle, callers cannot forget it
static <T> T withContext(RequestContext c, Supplier<T> body) {
CTX.set(c);
try { return body.get(); } finally { CTX.remove(); }
}
// Propagating to another pool requires cleanup on that thread too
Executor tracing = task -> {
RequestContext captured = CTX.get();
delegate.execute(() -> {
CTX.set(captured);
try { task.run(); } finally { CTX.remove(); }
});
};How to stop it coming back
- Wrap every application executor in a decorator that clears known holders after each task and fails the build if one is dirty
- Assert an empty context after each integration test rather than only before — leaks are visible on the way out, not the way in
- Run at least one CI suite with a single-threaded pool, which turns thread-reuse bleed from a race into a deterministic failure
- Keep context objects to identifiers and primitives; treat a `ThreadLocal` holding an entity or a bean as a review blocker
- Alarm on live instance counts of context types staying flat but non-zero while traffic is idle
- On Java 21+, adopt `ScopedValue` for new request-scoped state so the cleanup obligation does not exist
Practise this failure in a real repository
Gronex ships this as a runnable repository: pooled workers that populate a per-request holder, a failure path that skips the cleanup, and a test suite that asserts both the retained set stays bounded and that no task observes a predecessor’s context. Shrinking the payload reduces the leak without making the suite pass.
FAQ
ThreadLocalMap keys are weak references — why does anything leak?
The key is weak, the value is strong. Collecting the `ThreadLocal` object turns the entry stale but does not free the value; the value is reachable from the live `Thread`. Stale entries are cleared opportunistically during later `get`, `set`, or `remove` calls on that thread, so an idle pooled thread can hold a stale value indefinitely.
Is `set(null)` equivalent to `remove()`?
No. `set(null)` writes a null into a live entry, leaving the map slot and its weak key in place; `remove()` deletes the entry. Beyond the small retention difference, `set(null)` destroys your ability to distinguish "cleared" from "never set", which is exactly the distinction you need when debugging bleed.
Why does this never fail in tests?
Tests typically run one request per thread, or on a fresh thread per test, so no thread is ever reused across two requests with different contexts. The bug needs reuse. Pinning the test executor to one thread and asserting the holder is empty after each task exposes it immediately.
Do virtual threads make the problem go away?
They remove the retention half of it, because a virtual thread is not pooled — it is created per task and its thread-locals die with it. They do not remove the propagation problem across async boundaries, and `ThreadLocal` on millions of virtual threads costs memory per thread, which is why `ScopedValue` exists.
Can this cause a classloader leak rather than a heap leak?
Yes, and it is the classic cause of one. If the retained value is an instance of a class defined by a reloadable classloader, the live pool thread pins that loader, and the loader pins every class it defined plus their statics. The symptom is Metaspace growth across redeploys with a flat heap.
Related
Other errors engineers hit next to this one
- java.lang.OutOfMemoryError: Java heap space
- java.lang.OutOfMemoryError: Metaspace
- java.lang.OutOfMemoryError: GC overhead limit exceeded
- java.util.ConcurrentModificationException
- OutOfMemoryError: unable to create new native thread
- RejectedExecutionException: Task rejected from ThreadPoolExecutor
- FATAL: sorry, too many clients already
- Sessions stuck in "idle in transaction"