Concurrency
A ReentrantLock held forever after an exception — threads blocked with no deadlock reported
Written and reviewed by Sahil Srivastav
"http-nio-8080-exec-14" #58 prio=5 os_prio=0 tid=0x00007f2c1c0d9800 nid=0x5a1f waiting on condition [0x00007f2bd7bfe000]
java.lang.Thread.State: WAITING (parking)
at jdk.internal.misc.Unsafe.park(java.base@21/Native Method)
- parking to wait for <0x000000070f3c9a80> (a java.util.concurrent.locks.ReentrantLock$NonfairSync)
at java.util.concurrent.locks.LockSupport.park(java.base@21/LockSupport.java:221)
at java.util.concurrent.locks.AbstractQueuedSynchronizer.acquire(java.base@21/AbstractQueuedSynchronizer.java:754)
at java.util.concurrent.locks.ReentrantLock.lock(java.base@21/ReentrantLock.java:322)
at com.example.inventory.StockRegistry.reserve(StockRegistry.java:64)
Locked ownable synchronizers:
- NoneWhat this error actually means
There is no exception for this failure. The exception that caused it was thrown, logged, and handled somewhere upstream minutes earlier; what you are looking at now is the aftermath — a queue of threads parked in `AbstractQueuedSynchronizer.acquire` waiting for a lock whose owner has long since returned to the pool and is happily serving other requests.
The mechanism is the gap between `lock()` and `try`. `ReentrantLock` does not release on scope exit, on thread death, or on stack unwinding — it is an object with a counter and an owner field, and the only thing that decrements it is an `unlock()` call. If anything throws after the acquire and before the `finally` that releases, the counter stays above zero and the owner stays set to a thread that no longer knows it holds anything. Because reentrant locks track an owner rather than a stack frame, there is nothing in the runtime that can notice the mismatch.
The tell in the dump is the absence of a cycle. The JVM deadlock detector reports cycles in the wait-for graph, and a leaked lock is not a cycle — it is a single unreachable release. So `Found one Java-level deadlock` never appears, and the second giveaway is that the blocked threads name a lock whose *owner thread is nowhere in the dump waiting for anything*, or whose `Locked ownable synchronizers` section attributes it to a thread that is idle or gone.
This is also why the failure looks like a slow leak of capacity rather than an outage. One code path throws once per hour; each occurrence permanently removes the ability of any thread to enter that critical section. The endpoint that dies is not the endpoint that threw.
Causes, most common first
- 1`lock()` called before the `try` block. The dominant cause, and it looks correct at a glance because `unlock()` is in a `finally`. If the acquire itself sits above the `try`, any statement between the two — a log line that dereferences a null, a metric lookup, an argument validation — throws with the lock held and skips the `finally` entirely because no `try` was ever entered.
- 2Release placed at the end of the method body instead of in `finally`. Structurally the same bug with no `try` at all. It survives every test because tests exercise the path that reaches the last line. The first production exception turns it into a permanent outage of that code path.
- 3Early `return` or `break` inside the critical section. A guard clause added later — `if (stock.isEmpty()) return false;` — returns past the `unlock()`. Reviews miss this because the diff adds two lines and touches nothing to do with locking.
- 4Conditional acquire, unconditional release, or the reverse. `tryLock()` returns `false`, the code proceeds anyway, and the `finally` calls `unlock()` on a lock it never obtained. That throws `IllegalMonitorStateException`, which masks the real bug and, in a loop, can leave the counter permanently skewed when acquires and releases are unbalanced in a reentrant method.
- 5Reentrant acquire without a matching release count. `ReentrantLock` counts holds. A method that acquires twice — directly or via a helper that also locks — and releases once leaves the hold count at one forever. The owner thread itself notices nothing, since it can re-enter freely.
- 6Lock held across a blocking call that never returns. Not strictly a leak, but indistinguishable in the dump until you read the owner’s stack: an HTTP or JDBC call with no timeout inside the critical section. The lock is legitimately held, just for an unbounded duration. Check the owner’s frames before concluding the release was skipped.
When you see it
- One endpoint hangs indefinitely while every other endpoint is served normally
- CPU falls rather than rises — the stuck threads are parked, not spinning
- A restart resolves it fully, then it returns after roughly the same volume of failed requests
- Thread count grows steadily as new requests pile up behind the same critical section
- Grepping the dump for `Found one Java-level deadlock` returns nothing, which sends people looking in the wrong direction
- An earlier error in the logs — validation failure, downstream timeout, `NullPointerException` — precedes the first hang by minutes
How to diagnose it
Step 1
Dump with ownable synchronizers, not a plain dump
`ReentrantLock` is not a monitor, so a bare thread dump shows the waiters but attributes ownership poorly. `-l` adds the `Locked ownable synchronizers` section, which is what names the current owner of an AQS-based lock.
jcmd <pid> Thread.print -l > dump.txtStep 2
Find the lock identity, then find its owner
Take the hex identity from the `parking to wait for` line of any blocked thread and search the same dump for it in the ownable-synchronizers sections. If no thread lists it, the lock is held by a thread that has moved on — that is a leaked lock, conclusively.
grep -n "0x000000070f3c9a80" dump.txtStep 3
Prove the stacks are frozen, not slow
Two dumps several seconds apart. Identical frames on the waiters and a *different* stack on the owner between dumps means the owner is doing unrelated work while still holding the lock.
jcmd <pid> Thread.print -l > d1.txt; jcmd <pid> Thread.print -l > d2.txt; diff d1.txt d2.txtStep 4
Check whether the hold count can ever reach zero
In a development build, expose the lock’s state. A `getHoldCount()` of zero with `isLocked()` true means it is held by *another* thread; a non-zero count on an idle request thread is the smoking gun.
log.info("locked={} holdCount={} queued={}", lock.isLocked(), lock.getHoldCount(), lock.getQueueLength());Step 5
Reproduce by forcing the exception path
Call the operation once with input that throws, then call it normally from another thread. If the second call never returns, you have reproduced the bug in a unit test — and you now have a regression test that a reordered `lock()` would fail.
The fix
Adopt one acquire shape and never deviate: `lock()` on the line immediately before `try`, with nothing between them, and `unlock()` as the first statement of `finally`. Nothing between the acquire and the `try` is the whole point — that gap is the only place this bug can live.
For `tryLock`, branch on the result and release only inside the branch that acquired. The release must be governed by the same boolean that the acquire returned, never by reaching a particular line.
Remove early returns from critical sections, or restructure so the guard runs before the acquire. Computing the decision outside the lock and entering it only to mutate usually shortens the critical section as a side benefit, which helps contention too.
Never hold a lock across a network call, a disk write, or anything else that can block without a bound. If the data the call needs is under the lock, copy it out, release, call, then re-acquire to apply the result — and re-validate, because the state can have changed while you were outside.
Where the critical section is genuinely simple, delete the lock. A `ConcurrentHashMap.compute` or an `AtomicReference.updateAndGet` performs the read-modify-write atomically with no release to forget. Code that cannot leak a lock is better than code that releases correctly.
As a safety net once the shape is fixed, use `tryLock(timeout)` at the outermost acquisition so a held lock degrades into a failed request with an alert instead of an unbounded hang. Treat this as instrumentation, not as the fix.
// Leaks the lock: anything throwing before `try` skips the finally entirely
lock.lock();
validate(request); // throws -> lock held forever
try {
stock.put(sku, remaining);
} finally {
lock.unlock();
}
// Also leaks: the guard returns past the release
lock.lock();
try {
if (stock.isEmpty()) return false; // fine
...
} finally { lock.unlock(); }
// Correct: nothing between acquire and try; conditional acquire, conditional release
validate(request); // outside the critical section
if (!lock.tryLock(200, TimeUnit.MILLISECONDS)) {
throw new ReservationBusyException(sku);
}
try {
return stock.reserve(sku, qty);
} finally {
lock.unlock();
}How to stop it coming back
- Make "nothing between `lock()` and `try`" a review rule; it is mechanically checkable and catches every instance of this bug
- Enable an error-prone or SpotBugs rule for unreleased locks in CI rather than relying on human attention
- Write a test per critical section that forces the exception path and then asserts a subsequent call still completes within a timeout
- Alarm on `lock.getQueueLength()` sustained above zero — waiters queuing while throughput is normal is the earliest signal
- Prefer atomic collection operations over explicit locks wherever the critical section is a single read-modify-write
- Capture a thread dump automatically when a request exceeds a latency ceiling, so the evidence exists without needing to catch it live
Practise this failure in a real repository
Gronex ships this as a runnable repository: a registry whose acquire sits outside the `try`, an input that throws on one branch, and a test suite that drives the failing path and then asserts the next caller still makes progress inside a deadline. Adding a timeout makes the hang visible but does not make the suite pass.
FAQ
Why is no deadlock reported when threads are clearly stuck on a lock?
The JVM detector looks for cycles in the wait-for graph. A leaked lock is not a cycle — one owner simply never releases, and the owner is not itself waiting for anything. Absence of `Found one Java-level deadlock` rules out a cycle; it says nothing about whether a lock is stuck.
Does `synchronized` have the same problem?
No, and that is its one real advantage. A `synchronized` block releases the monitor when the frame exits for any reason, including an exception, because the release is emitted by the compiler as an exception-table handler. You trade that safety for `tryLock`, timeouts, fairness, and multiple conditions when you move to `ReentrantLock` — so you take on the release obligation.
Is the lock released when the owning thread dies?
Not for `ReentrantLock`. The lock holds a reference to the owner thread and a hold count; neither is cleared by thread termination, so the lock stays held permanently and every waiter parks forever. A monitor taken with `synchronized` is released during unwinding, which is why thread death is survivable there and not here.
Would raising the number of worker threads help?
It delays the visible outage and makes the eventual dump harder to read. Every new thread reaching that critical section parks, so you are adding threads to a queue that can never drain. Capacity is not the constraint — a release that never happens is.
How is this different from a lock held across a slow call?
The waiters look identical; the owner does not. Read the owner’s stack. Frames inside a socket read or a JDBC call mean the lock is legitimately held for too long, and the fix is to move the call out of the critical section. An owner doing something unrelated, or no listed owner at all, means the release was skipped.
Related
Other errors engineers hit next to this one
- ERR_STREAM_PREMATURE_CLOSE during an upload
- Process exits before asynchronous writes finish
- ERR_MODULE_NOT_FOUND during ESM/CommonJS migration
- Unbounded JSON body blocks the loop or exhausts memory
- Undici / fetch connections remain occupied
- ERR_UNHANDLED_REJECTION in a worker thread
- command not found in a script that works interactively
- Permission denied when executing a script