Concurrency

Thread pool starvation with nested tasks

Written and reviewed by Sahil Srivastav

ConcurrencyExecutorsLiveness
"batch-worker-1" #24 prio=5 os_prio=0 tid=0x00007fa41c0b2800 nid=0x3f01 waiting on condition [0x00007fa3f4bfd000]
   java.lang.Thread.State: WAITING (parking)
	at jdk.internal.misc.Unsafe.park(java.base@21/Native Method)
	at java.util.concurrent.locks.LockSupport.park(java.base@21/LockSupport.java:221)
	at java.util.concurrent.FutureTask.awaitDone(java.base@21/FutureTask.java:500)
	at java.util.concurrent.FutureTask.get(java.base@21/FutureTask.java:190)
	at com.example.batch.OrderEnricher.enrich(OrderEnricher.java:73)
	at java.util.concurrent.ThreadPoolExecutor.runWorker(java.base@21/ThreadPoolExecutor.java:1144)

"batch-worker-2" ... identical stack
"batch-worker-3" ... identical stack
"batch-worker-4" ... identical stack

[monitor] pool=batch active=4 max=4 queue=312 completed=0 (last 5m)

What this error actually means

Every worker in the pool is occupying a thread while waiting for work that can only run on a thread from that same pool. The queue holds the sub-tasks, the sub-tasks need a worker, and every worker is blocked waiting for a sub-task. Nothing will break it and nothing will time out, because nobody is waiting on a lock — they are waiting on a `Future` that is queued behind them.

This is a deadlock, but it is invisible to the JVM deadlock detector. The detector finds cycles in the *lock* wait-for graph, and there is no lock here: the cycle runs through a queue and a scheduler, which the JVM does not model. So `Found one Java-level deadlock` never appears, and the incident presents as a silent, complete loss of throughput with clean error metrics.

The reason it survives review is that the dependency is not visible at the call site. A handler calls a helper, the helper fans out to an executor and joins, and the fact that the handler is *itself* running on that executor is three layers away in a different file. The bug is a property of the composition, not of either piece — which is why it appears after an unrelated refactor moves a caller onto a pooled thread.

Thread count is what determines whether you see it at all. With `n` workers and a task that spawns one sub-task and waits, the pool survives until `n` such tasks run concurrently. A pool of 32 hides the bug under normal load and fails at peak; a pool of 2 fails immediately. That is why "it only happens in production" is the usual framing, and why shrinking the pool is the fastest way to reproduce it.

Causes, most common first

  1. 1A pooled task submits to its own pool and joins. The direct form. The outer task holds a worker for the whole duration of the inner task, and the inner task cannot start until a worker is free. With enough concurrent outer tasks, forward progress requires a worker that does not exist.
  2. 2A parallel stream inside work already on the common ForkJoinPool. Parallel streams use `ForkJoinPool.commonPool()` unless you submit them into your own pool. Nesting one inside another piece of common-pool work, or inside a `CompletableFuture` stage that also defaults to the common pool, contends for the same small, JVM-wide set of workers — whose default parallelism is one less than the available processor count.
  3. 3A `CompletableFuture` chain that joins on the default executor. `supplyAsync` and the non-`Async` continuations run on the common pool by default, so a `join()` inside a stage blocks a common-pool worker while the stage it awaits waits for one. Mixing library code that also uses the common pool makes the contention cross-cutting and very hard to attribute.
  4. 4Fan-out with a scatter-gather barrier. The outer task splits work into `k` sub-tasks and waits for all of them. Each outer task now needs `k + 1` simultaneous slots. The pool size that works for one request fails for two, and the failure threshold depends on `k`, which is data-dependent — so it triggers on one unusually large input.
  5. 5A shared pool crossing layers that do not know about each other. An application-wide executor injected everywhere. Layer A submits and joins; layer B, unaware, also submits. The nesting emerges from the wiring rather than from any single piece of code, so nothing in review looks wrong.
  6. 6A bounded resource acquired before the submission. The same shape with a different scarce thing: the outer task holds a database connection or a semaphore permit while waiting for a sub-task that also needs one. Sizing the pool correctly does not help, because the constraint is the other resource.

When you see it

  • Throughput goes to zero with no exceptions logged and no error rate change
  • Every thread in one pool has an identical stack ending in `FutureTask.get` or `CompletableFuture.join`
  • Queue depth climbs steadily while completed-task count stays frozen
  • CPU drops to idle — the workers are parked, not spinning
  • No deadlock is reported by the JVM, which misdirects the investigation towards the database or a downstream service
  • It clears on restart and reappears at the same load level, often at exactly the same time each day
  • Reducing the pool size makes it reproduce instantly, which is the opposite of what a capacity problem does

How to diagnose it

Step 1

Take a dump and look for identical stacks across one pool

This is diagnostic on its own. `n` threads from the same named pool, all parked in `FutureTask.awaitDone` or `CompletableFuture` waiting code, with your application frame just below, is starvation and nothing else. A capacity problem shows threads doing *different* things.

jcmd <pid> Thread.print | grep -A 10 "batch-worker" | grep -c "FutureTask.awaitDone"

Step 2

Read active, queue depth, and completed together

Active at maximum, queue non-empty, completed not advancing is the fingerprint. Under genuine saturation, completed advances — slowly, but it advances. Frozen completion count with a full pool proves no work is finishing at all.

Metrics.gauge("pool.active", pool, ThreadPoolExecutor::getActiveCount);
Metrics.gauge("pool.queue", pool, p -> p.getQueue().size());
Metrics.gauge("pool.completed", pool, ThreadPoolExecutor::getCompletedTaskCount);

Step 3

Reproduce by shrinking the pool to one

The inverted test that settles the argument. Nested submission with a single worker deadlocks on the first request, deterministically. If a one-thread pool hangs where a larger one is merely slow, the dependency is structural.

var pool = Executors.newFixedThreadPool(1);  // nested join hangs immediately

Step 4

Find the nesting statically

Grep for joins and blocking gets, then check each one against which pool the enclosing code runs on. The answer is often surprising, because the enclosing pool is decided by a caller in another module.

grep -rn -E "\\.get\\(\\)|\\.join\\(\\)|invokeAll|allOf\\(.*\\)\\.join" src/main/java | grep -v test

Step 5

Name every pool before you need to

Default `pool-3-thread-7` names make it impossible to tell whether the identical stacks belong to one pool or three. A thread factory with a prefix per pool turns a twenty-minute dump-reading exercise into a `grep -c`.

Executors.newFixedThreadPool(8, Thread.ofPlatform().name("batch-worker-", 1).factory());

The fix

Remove the join rather than resize the pool. Restructure the fan-out so the outer task submits the sub-tasks and *returns*, with the aggregation done by a continuation that runs when the children complete. `CompletableFuture.allOf(...).thenApply(...)` expresses the join without any thread waiting for it — nothing occupies a worker while the children queue, so the cycle cannot form regardless of pool size or fan-out width.

If the code must block, give the blocking wait a different pool from the work it waits for. Two pools with a strict one-directional dependency — outer waits on inner, inner never submits to outer — cannot starve, because the inner pool’s workers are never held by a waiter. Make the direction explicit in the wiring and enforce it, or you will reintroduce the cycle in six months.

Never run a parallel stream, or a defaulted `CompletableFuture` stage, inside work that is itself on the common ForkJoinPool. Submit the stream into a dedicated `ForkJoinPool` you own, and always pass an explicit executor to `supplyAsync`, `thenApplyAsync`, and friends. Relying on the default executor is how unrelated subsystems end up sharing one small pool.

Make fan-out width bounded and independent of input size. A task that spawns one sub-task per row needs a slot count proportional to the data; chunking the work into a fixed number of batches makes the concurrency requirement a constant you can reason about.

Where a scarce resource is held across the wait — a connection, a permit — release it before submitting and re-acquire after. Holding a pool slot *and* a connection while waiting for a task that needs both is two overlapping cycles, and fixing only the pool leaves the other.

Add a bounded `get(timeout, unit)` at every remaining blocking join once the structure is right. This does not fix starvation, but it converts an eternal silent hang into a timeout you can alarm on — the difference between a five-minute incident and a five-hour one.

// Deadlocks once `n` outer tasks run at the same time: each holds a worker
// while its sub-task sits in the queue behind it.
var pool = Executors.newFixedThreadPool(4);

Future<Report> outer = pool.submit(() -> {
    Future<Rows> inner = pool.submit(this::loadRows);   // same pool
    return render(inner.get());                         // holds a worker, waits
});

// Fix 1 (preferred): no thread ever waits — the join becomes a continuation
CompletableFuture<Report> outer =
    CompletableFuture.supplyAsync(this::loadRows, ioPool)
                     .thenApplyAsync(this::render, cpuPool);

// Fix 2: blocking is fine if the waiter and the worker are different pools,
// with a strictly one-directional dependency.
var orchestrator = Executors.newFixedThreadPool(4, named("orchestrator-"));
var workers      = Executors.newFixedThreadPool(8, named("worker-"));

orchestrator.submit(() -> {
    var parts = workers.invokeAll(chunks(order));   // workers never submit back
    return merge(parts);
});

How to stop it coming back

  • Make it an architectural rule that a task never blocks on a task in its own pool, and record for every pool which pools it is allowed to depend on
  • Always pass an explicit executor to `CompletableFuture` async stages; treat a defaulted stage in application code as a review finding
  • Name every pool through a thread factory — diagnosis time for this failure is dominated by whether the threads are identifiable
  • Run a CI suite with every pool sized to one thread: structural nesting hangs immediately, and a timeout in the test is a clear failure signal
  • Alarm on completed-task count being flat while queue depth is non-zero, which detects starvation without needing a dump
  • Keep fan-out width constant rather than proportional to input, so the concurrency requirement does not depend on the data

Practise this failure in a real repository

Gronex ships this as a runnable repository: pooled tasks that fan out into their own bounded executor and wait for the results. The test suite drives concurrent requests and asserts forward progress inside a deadline, so enlarging the pool moves the threshold without passing — only removing the dependency does.

FAQ

Why is no deadlock reported when the pool is clearly deadlocked?

The JVM detector only finds cycles among monitors and AQS-based locks. Here the cycle runs through a work queue and a scheduling policy, which the JVM has no model of. The threads are parked on a `Future`, which is legitimate waiting as far as the runtime is concerned, so nothing is reported.

Would a larger pool fix it?

It raises the concurrency level at which the pool fails and nothing else. With `n` workers and one nested wait per task, `n` concurrent tasks still exhausts it, and wider fan-out lowers the threshold further. The failure depends on the dependency existing, not on the pool being small.

Does ForkJoinPool avoid this?

Partly, and not reliably enough to depend on. A worker that calls `join()` on a ForkJoinPool task can help execute other queued tasks rather than idling, and `ManagedBlocker` lets the pool compensate for a genuinely blocking wait by starting another worker. But a plain blocking call — a socket read, a `Future.get` on a different executor — gets neither benefit, and the common pool’s default parallelism of processors minus one is small enough that a couple of blocked workers matter.

How do I tell starvation from ordinary saturation?

Watch the completed-task counter. Saturation is slow progress: completed rises, latency is bad, throughput is non-zero. Starvation is no progress: completed is frozen while active sits at maximum. The stacks say the same thing — different frames under saturation, identical waiting frames under starvation.

Do virtual threads solve it?

They remove the scarcity that causes it, which is not the same as making the design sound. With a virtual thread per task there is no fixed worker count to exhaust, so a blocking join costs memory instead of liveness. The cycle can still form around any other bounded resource — a connection pool, a semaphore — so the "never wait on work that needs a resource you are holding" rule still applies.

Related

Other errors engineers hit next to this one

Full error and symptom index →