Concurrency

Found one Java-level deadlock

Written and reviewed by Sahil Srivastav

ConcurrencyLocksProduction incident
Found one Java-level deadlock:
=============================
"transfer-worker-3":
  waiting to lock monitor 0x00007f8c0400a1b8 (object 0x000000076ab2c108, a com.example.Account),
  which is held by "transfer-worker-7"
"transfer-worker-7":
  waiting to lock monitor 0x00007f8c04009c90 (object 0x000000076ab2c0e8, a com.example.Account),
  which is held by "transfer-worker-3"

What this error actually means

The JVM has done the hard part. This block appears at the bottom of a thread dump when the deadlock detector finds a cycle in the lock wait-for graph: thread A holds lock 1 and waits for lock 2 while thread B holds lock 2 and waits for lock 1. Neither will ever proceed, and nothing will break the cycle short of killing the process.

Read it as a cycle, not as a list. Each stanza names the monitor a thread is waiting for and the thread that holds it. Follow the "held by" links and you return to where you started — that loop is the bug, and the object identities in it tell you exactly which two resources were locked in opposite orders.

The structural cause is almost always the same: two code paths acquire the same pair of locks in different orders. A transfer from account X to Y locks X then Y; a simultaneous transfer from Y to X locks Y then X. The code looks symmetric and correct, and it deadlocks only when both directions run at once — which is why it passes review and fails in production.

Causes, most common first

  1. 1Inconsistent lock ordering across two paths. The dominant cause. Two operations lock the same two objects in opposite orders because each is written from its own point of view. Deadlock requires only that both run concurrently with mirrored arguments.
  2. 2Nested synchronized blocks with a callback in between. Holding a lock while invoking code you do not control — a listener, an override, an injected strategy — means that code may acquire another lock. The second acquisition is invisible at the call site, so the ordering violation is not local.
  3. 3Lock acquired inside a database transaction, or the reverse. Mixing application monitors with database row locks creates a cycle spanning two systems. The JVM detector only sees its own half, so the dump may show threads blocked with no Java-level cycle reported at all.
  4. 4Locking on mutable or interned objects. Synchronising on a `String` literal, a boxed `Integer`, or a field that gets reassigned means the identity of the monitor is not what you think. Two unrelated components can end up sharing a monitor, producing cycles across code that was never meant to interact.
  5. 5Lock not released on the exception path. With explicit `Lock` objects, an exception between `lock()` and `unlock()` leaves the lock held forever. Every subsequent acquirer blocks. The dump shows many threads waiting on a lock whose owner is doing something unrelated — or has already finished.

When you see it

  • A subset of requests hang forever while the rest are served normally
  • CPU drops rather than rises, because the stuck threads are parked, not spinning
  • Thread count climbs as new requests queue behind the same locks
  • It reproduces only under concurrency and only with the specific pair of entities involved
  • Health checks pass if they do not touch the affected path, so the instance is never replaced

How to diagnose it

Step 1

Take a thread dump and read the bottom

The detector output is appended after the thread list. Two dumps a few seconds apart confirm it is a true deadlock rather than slow progress: in a deadlock the stacks are byte-for-byte identical.

jcmd <pid> Thread.print > dump1.txt; jcmd <pid> Thread.print > dump2.txt; diff dump1.txt dump2.txt

Step 2

Identify the two lock sites from the object types

The dump names the monitor’s object class and the exact line in each thread’s stack. Those two lines are the two acquisition sites. Compare the order in which each path takes the pair.

Step 3

Check for a held-but-not-owned pattern

If threads are blocked but no cycle is reported, suspect a lock leaked on an exception path or a cycle through the database. `jcmd Thread.print -l` includes ownable synchronizers, which covers `ReentrantLock` where plain dumps show less.

jcmd <pid> Thread.print -l | grep -B 5 "waiting on condition"

Step 4

Reproduce with mirrored arguments

A test that runs the operation in both directions concurrently — X to Y and Y to X, on many threads — reproduces ordering deadlocks reliably. Single-direction load tests never will, which is why this class of bug ships.

The fix

Impose a global lock order and make every path follow it. Derive the order from a stable, total ordering of the resources — the primary key, or `System.identityHashCode` with an explicit tie-breaker — and always acquire in that order regardless of which direction the operation logically runs. Once every path agrees on the order, a cycle is impossible by construction. This is the fix to prefer, because it removes the possibility rather than handling the symptom.

Shrink the critical section so the pair is never held at once. Often the second lock is needed only for a brief mutation that can be done after the first is released, or the whole operation can be restructured around a single lock covering both resources.

Never hold a lock while calling code you do not control. Compute inside the lock, release, then invoke callbacks and listeners outside it. This eliminates the entire class of invisible nested acquisition.

Use `tryLock` with a timeout as a safety net, not a design. Acquiring with a bound turns an eternal hang into a failed operation you can retry with backoff and, more importantly, an alert. It is the right belt-and-braces addition after ordering is fixed, and a poor substitute for fixing it.

Always release in `finally`. For explicit locks, `lock()` immediately before `try` and `unlock()` in `finally` is the only safe shape — this also fixes the leaked-lock variant where no cycle exists at all.

// Deadlocks when transfer(a,b) and transfer(b,a) run concurrently
void transfer(Account from, Account to, BigDecimal amt) {
    synchronized (from) {
        synchronized (to) { from.debit(amt); to.credit(amt); }
    }
}

// Total order on identity makes a cycle impossible
void transfer(Account from, Account to, BigDecimal amt) {
    Account first  = from.id() < to.id() ? from : to;
    Account second = first == from ? to : from;
    synchronized (first) {
        synchronized (second) { from.debit(amt); to.credit(amt); }
    }
}

How to stop it coming back

  • Document the lock order for every pair of locks that can be held together, and enforce it in review
  • Prefer a single coarser lock, or lock-free structures, over two fine-grained locks whose interaction you must reason about
  • Write concurrency tests with mirrored arguments across many threads; that is the only shape that catches ordering bugs
  • Alarm on requests exceeding a latency ceiling and dump threads automatically when it trips, so the evidence is captured while the deadlock is live
  • Never synchronise on interned strings, boxed primitives, or reassignable fields

Practise this failure in a real repository

Gronex ships this as a runnable repository: a transfer path that deadlocks under mirrored concurrent calls, with a test suite that drives both directions hard and asserts progress plus ledger consistency. Adding a timeout does not make it pass.

FAQ

Why does the JVM detect it but not recover?

Deadlock detection is diagnostic only. Breaking a cycle means forcibly taking a lock from a thread mid-critical-section, which would leave shared state arbitrarily corrupted. Reporting and leaving the process intact is the safer choice, and it is why you must design the cycle out.

Does using ReentrantLock instead of synchronized help?

Only because it offers `tryLock` with a timeout, which converts a hang into a recoverable failure. It does not remove the ordering bug. Reentrancy protects you from self-deadlock on the same lock, never from a cycle across two.

Why did this never happen in testing?

Because it requires both orderings to be in flight simultaneously on the same pair of resources. Sequential tests and single-direction load tests cannot produce that interleaving. You need concurrent mirrored arguments, which is a test people write only after the first incident.

What if threads are blocked but no deadlock is reported?

Then the cycle is not entirely inside the JVM. Likely candidates: a lock leaked on an exception path so the owner never releases it, a cycle through database row locks, or a `ReentrantLock` held across a blocking I/O call with no timeout.

Related

Other errors engineers hit next to this one

Full error and symptom index →