Concurrency
ExecutorService shutdown hangs in awaitTermination
Written and reviewed by Sahil Srivastav
2026-04-02T18:22:10.114Z INFO Lifecycle - shutting down executor batch-pool
2026-04-02T18:22:40.118Z WARN Lifecycle - executor did not terminate in 30s, forcing
2026-04-02T18:23:10.121Z WARN Lifecycle - executor still has 3 active tasks after shutdownNow()
# process does not exit; SIGTERM already delivered 60s ago; orchestrator will SIGKILL
"batch-worker-3" #29 prio=5 os_prio=0 tid=0x00007f1a2c0c1000 nid=0x4a13 runnable [0x00007f19d8ffd000]
java.lang.Thread.State: RUNNABLE
at java.base/sun.nio.ch.Net.poll(Native Method)
at java.base/sun.nio.ch.NioSocketImpl.park(NioSocketImpl.java:186)
at java.base/sun.nio.ch.NioSocketImpl.timedRead(NioSocketImpl.java:285)
at com.example.feed.FeedClient.read(FeedClient.java:88)What this error actually means
`shutdown()` is a request to stop accepting new work, not a request to stop current work. It marks the executor as shutting down, refuses further submissions with `RejectedExecutionException`, and then lets every already-submitted task — running *and queued* — run to completion. `awaitTermination` returns only when all of them have finished. So a pool with a thousand queued tasks will not terminate until it has executed all thousand, no matter how long that takes.
`shutdownNow()` is stronger but still not a kill. It drains and returns the queued tasks, then calls `Thread.interrupt()` on each worker. Interruption is cooperative: it sets a flag and, for a thread parked in a *interruptible* operation, causes that operation to throw. A thread that is not in such an operation, or that catches `InterruptedException` and carries on, is unaffected. There is no mechanism in Java to forcibly stop a thread — `Thread.stop` was deprecated and then removed because it unlocked monitors at arbitrary points and left shared state corrupted.
The practical consequence is the trap in the dump above: a task blocked reading from a plain `Socket` shows as `RUNNABLE`, and interrupting it does nothing. Operations on `InterruptibleChannel` respond to interrupt by closing the channel and throwing `ClosedByInterruptException`; classic socket I/O, most JDBC drivers, and native calls do not. The only thing that unblocks them is a read timeout or someone closing the underlying resource.
On top of all this, pool threads created by `Executors` factories are **non-daemon**, so a JVM that has finished `main` will not exit while one survives. The visible failure is a container that ignores SIGTERM, sits out the grace period, and is SIGKILLed — losing in-flight work, skipping cleanup, and producing the deployment that takes thirty seconds longer per pod than it should.
Causes, most common first
- 1A task blocked in an uninterruptible I/O call. The dominant cause. A socket read, a JDBC statement, or a native call with no timeout. Interrupt sets a flag the call never checks. Nothing short of a timeout at the protocol level or closing the socket will unblock it, which is why "add a timeout to every outbound call" is a shutdown fix as much as a latency fix.
- 2A long queue that `shutdown()` insists on draining. Thousands of queued tasks each taking a second. `shutdown()` is behaving exactly as documented; the design error is expecting graceful and immediate from the same call. Use `shutdownNow()` to drain the queue, and re-enqueue the returned tasks durably if they matter.
- 3Tasks that catch `InterruptedException` and continue. A `while (true)` loop whose body catches the exception, logs it, and loops again. Because the exception *clears* the interrupt flag, even a well-behaved check later in the loop sees a non-interrupted thread. The task is now unstoppable by any means the platform provides.
- 4Non-daemon pool threads outliving `main`. Default `Executors` thread factories create non-daemon threads, so one surviving worker keeps the JVM alive with no error and no log line. Common with a pool created by a library or a static initialiser that no shutdown hook knows about.
- 5Shutdown ordering that lets producers keep submitting. The pool is shut down while upstream components still submit. Every submission is rejected, error handling retries, and the shutdown is racing a producer. The visible symptom is a burst of `RejectedExecutionException` during every deploy.
- 6Periodic tasks on a scheduled executor. A `ScheduledThreadPoolExecutor` with a fixed-rate task will not terminate while the task remains scheduled. `shutdown()` cancels pending periodic tasks by default only if the policy says so; the currently executing run still has to finish, and a task whose body blocks indefinitely holds termination open.
- 7`awaitTermination` called on the wrong thread or not at all. Calling it from a shutdown hook that itself runs on a thread the pool is waiting for, or awaiting one pool while a second still feeds it. A dependency cycle among pools at shutdown is the same starvation problem in a different costume.
When you see it
- `awaitTermination` returns `false` at every timeout while the same tasks stay active
- The process survives SIGTERM and is killed by the orchestrator after the grace period
- Rolling deploys are slow by exactly the termination grace period for every instance
- Logs show "forcing" and then nothing, because `shutdownNow` returned and the workers ignored it
- A thread dump shows workers `RUNNABLE` inside `sun.nio.ch` or a JDBC driver rather than parked
- Shutdown succeeds on an idle instance and hangs on a busy one, so it never reproduces in staging
- `RejectedExecutionException` appears from other components during shutdown because the pool closed while producers were still running
How to diagnose it
Step 1
Dump during the hang and read the worker states
This immediately splits the cause in two. `RUNNABLE` inside `sun.nio.ch` or a driver means uninterruptible I/O — the fix is a timeout or closing the resource. `WAITING` or `TIMED_WAITING` in your own code means the task is ignoring or never checking interruption.
jcmd <pid> Thread.print | grep -A 8 "batch-worker" | grep -E "Thread.State|sun.nio.ch|java.net|Driver"Step 2
Distinguish queued from running
If the queue is long, `shutdown()` is doing its documented job and no amount of waiting is a bug — it is a design mismatch. Log all three numbers at the start of shutdown so the post-mortem does not need a guess.
log.info("shutdown: active={} queued={} completed={}",
pool.getActiveCount(), pool.getQueue().size(), pool.getCompletedTaskCount());Step 3
Find the thread keeping the JVM alive
When the process will not exit after shutdown completes, list non-daemon threads. Anything other than your own known long-lived threads is a pool nobody is closing — often created inside a library.
jcmd <pid> Thread.print | grep -B 1 "os_prio" | grep -v daemon | head -30Step 4
Prove whether interruption is observed at all
Instrument the task loop to log the interrupt flag. A task that never logs a true value under `shutdownNow` is either blocked uninterruptibly or swallowing the exception, and the log tells you which.
log.info("iteration interrupted={}", Thread.currentThread().isInterrupted());Step 5
Test shutdown under load, not at rest
Drive the service at realistic concurrency, send SIGTERM, and assert the process exits well inside the grace period. Shutdown bugs are load-dependent by nature, so an idle shutdown test proves nothing.
kill -TERM <pid>; time wait <pid>The fix
Shut down in the right order: stop accepting external work, stop producers, then `shutdown()` the pool, `awaitTermination` with a bounded timeout, `shutdownNow()` if that expires, and await once more before giving up. Skipping the producer step is what generates the rejection storm; skipping the second await is what leaves you unsure whether force actually worked.
Give every outbound call a timeout. This is the fix for the uninterruptible-I/O case and there is no substitute: a connect timeout, a read timeout, a JDBC `queryTimeout` or statement timeout, and a client-level deadline. A task that cannot block forever cannot hold shutdown open forever, which bounds your termination time by construction.
Make tasks cancellation-aware. Check `Thread.currentThread().isInterrupted()` at the top of each loop iteration, and on catching `InterruptedException` either exit the loop or restore the flag with `Thread.currentThread().interrupt()` before returning. A task that neither exits nor restores the flag is unstoppable.
Close the resource to unblock what interruption cannot reach. Keep a reference to the socket or connection the task is using and close it during forced shutdown: the blocked read then fails with an `IOException` you can handle. This is the only reliable way to stop legacy blocking I/O.
For scheduled executors, cancel periodic tasks explicitly before shutdown and set `setRemoveOnCancelPolicy(true)` so cancelled tasks leave the queue rather than accumulating in it. Relying on shutdown to reap a fixed-rate task is fragile.
Make pool threads daemon threads *only* as a backstop for exit, never as a substitute for shutdown. Daemon threads are killed abruptly at JVM exit with no cleanup, so work in flight is lost silently — acceptable for a metrics reporter, unacceptable for anything writing data.
Match the timeouts to the platform’s grace period. If the orchestrator allows 30 seconds, budget graceful drain plus forced shutdown plus final await inside that, and derive the numbers from your slowest legitimate task rather than choosing them by feel. Two 30-second awaits in a 30-second grace period guarantee a SIGKILL.
// Hangs: producers still submitting, no force, no second await
pool.shutdown();
pool.awaitTermination(Long.MAX_VALUE, TimeUnit.NANOSECONDS);
// Reliable shutdown inside a bounded grace period
acceptingRequests = false; // 1. stop external intake
producers.stop(); // 2. stop anything that submits
pool.shutdown(); // 3. no new tasks; drain what is queued
try {
if (!pool.awaitTermination(15, TimeUnit.SECONDS)) {
List<Runnable> dropped = pool.shutdownNow(); // 4. interrupt + drain
log.warn("forcing shutdown, {} tasks dropped", dropped.size());
closeOpenSockets(); // 5. unblock uninterruptible I/O
if (!pool.awaitTermination(10, TimeUnit.SECONDS)) {
log.error("executor did not terminate; workers may be stuck in native I/O");
}
}
} catch (InterruptedException e) {
pool.shutdownNow();
Thread.currentThread().interrupt(); // never swallow it
}
// A cancellation-aware task loop
while (!Thread.currentThread().isInterrupted()) {
try {
Item item = queue.poll(1, TimeUnit.SECONDS);
if (item != null) process(item);
} catch (InterruptedException e) {
Thread.currentThread().interrupt(); // restore, then exit
break;
}
}How to stop it coming back
- Own every executor in a lifecycle component with an explicit close, and fail the build if a pool is created without one
- Set connect, read, and statement timeouts on every client library; treat an untimed outbound call as a shutdown defect as well as a latency one
- Include a SIGTERM-under-load test in CI that asserts exit well inside the production grace period
- Derive shutdown timeouts from the orchestrator’s grace period, and log when a phase exceeds its budget
- Use daemon threads only for work that is safe to lose abruptly, and document that choice at the thread factory
- Alarm on pods being SIGKILLed after the grace period — it is usually read as a deploy quirk when it is an unterminated pool
FAQ
What is the difference between `shutdown()` and `shutdownNow()`?
`shutdown()` stops new submissions and lets both running and queued tasks finish. `shutdownNow()` additionally discards the queue — returning those tasks to you — and interrupts the running workers. Neither forcibly stops anything: interruption is a cooperative request that a task can ignore.
Why does `shutdownNow()` not stop a task blocked on a socket read?
Interruption only unblocks operations that are documented as interruptible. `InterruptibleChannel` operations throw `ClosedByInterruptException`; classic `Socket` I/O, most JDBC calls, and native methods simply do not check the flag. The thread even shows as `RUNNABLE` in the dump. Only a timeout or closing the underlying socket will end it.
Why will the JVM not exit even after shutdown returns?
Almost always a surviving non-daemon thread. `Executors` factories create non-daemon threads by default, and a pool created by a library or a static initialiser may have no shutdown path at all. List non-daemon threads in a dump; whatever is there beyond your known set is what is holding the process open.
Should I make pool threads daemon threads to fix this?
Only as a last-resort backstop. Daemon threads are terminated at JVM exit with no unwinding and no cleanup, so in-flight work vanishes silently — fine for a metrics reporter, unacceptable for anything that writes data. It converts a visible hang into invisible data loss, which is usually the worse trade.
Is `Thread.stop()` an option for a task that will not die?
No. It was deprecated for releasing monitors at arbitrary points and leaving shared state in an inconsistent state, and it has since been made to throw unconditionally. If a task cannot be stopped cooperatively, the honest options are to bound its blocking with timeouts, close the resource it is blocked on, or let the process be replaced.
Related
Other errors engineers hit next to this one
- ERROR: deadlock detected
- ERROR: canceling statement due to statement timeout
- ERROR: could not serialize access due to concurrent update
- ERROR: current transaction is aborted, commands ignored until end of transaction block
- ERROR: duplicate key value violates unique constraint
- Deep OFFSET pagination getting slower every page
- ERROR: canceling statement due to lock timeout (ALTER TABLE)
- ERROR: canceling statement due to conflict with recovery