Concurrency

ReadWriteLock writer starvation

Written and reviewed by Sahil Srivastav

ConcurrencyLocksLiveness
"config-refresher" #42 prio=5 os_prio=0 tid=0x00007f9a4c0d3000 nid=0x6b17 waiting on condition [0x00007f99ec3fd000]
   java.lang.Thread.State: WAITING (parking)
	at jdk.internal.misc.Unsafe.park(java.base@21/Native Method)
	- parking to wait for  <0x00000006e2110c48> (a java.util.concurrent.locks.ReentrantReadWriteLock$NonfairSync)
	at java.util.concurrent.locks.ReentrantReadWriteLock$WriteLock.lock(java.base@21/ReentrantReadWriteLock.java:959)
	at com.example.config.ConfigStore.publish(ConfigStore.java:71)

[monitor] rwlock readLockCount=38 writeLocked=false queuedWriters=1 writerWaitMs=412900
# The writer has been waiting 6m52s. Readers are never all absent at the same instant.

What this error actually means

A write lock requires that *no* reader holds the lock. With a high enough read arrival rate, that condition can simply never be true: reader A is still inside when reader B arrives, B is still inside when C arrives, and the read count never reaches zero. The writer is queued, correctly, for a moment that never comes.

Whether new readers are allowed to overtake a queued writer depends on the lock’s policy, and the default is the permissive one. A non-fair `ReentrantReadWriteLock` lets an arriving reader acquire immediately if the lock is currently read-held, even with a writer waiting — barging is what makes the non-fair lock fast, and it is exactly what starves the writer. The documentation is explicit that non-fair mode makes no guarantees about ordering.

The practical shape is a background writer — a configuration refresher, a cache invalidator, a compaction task — that stops making progress under load and resumes when traffic drops. Nothing errors. The system serves stale data indefinitely, and the incident is usually discovered as "the config change never took effect on three instances", not as a concurrency problem.

Sitting next to this is a second, deterministic trap: the read lock cannot be upgraded. Calling `writeLock().lock()` while holding the read lock blocks forever, because the writer waits for all readers to leave and the caller is one of the readers. It self-deadlocks on the first execution, and the JVM does not report it as a deadlock because there is no cycle between two threads — one thread is waiting for itself. Downgrading is legal and supported: acquire the write lock, then acquire the read lock, then release the write lock.

Causes, most common first

  1. 1Non-fair mode with a continuous read stream. The default construction, `new ReentrantReadWriteLock()`, is non-fair. Arriving readers barge past the queued writer while the lock is read-held, so the read count never drains. This is the cause in the large majority of cases and it is a policy choice, not a bug in the lock.
  2. 2Long read critical sections. Read hold time is what determines whether the count can reach zero. A read section that does I/O, serialises a response, or iterates a large structure keeps overlapping with the next reader. Short reads starve writers far less, because gaps appear naturally.
  3. 3Attempting to upgrade from read to write. Holding the read lock and acquiring the write lock blocks permanently — the thread is waiting for itself to release. Deterministic, reproducible on the first call, and not reported as a deadlock because the cycle is within one thread.
  4. 4Reentrant reads nested in a call chain. A read lock taken in an outer method and again in a helper. Legal, since read locks are reentrant, but it doubles the effective hold and makes the zero-reader window correspondingly rarer. It also makes upgrade attempts more likely, because the outer read hold is invisible at the inner call site.
  5. 5A read lock leaked on an exception path. A read hold that is never released keeps the count permanently above zero, so the writer can never acquire. Indistinguishable from starvation in a gauge; distinguishable in a dump, where the read count is stable rather than fluctuating.
  6. 6Using a read-write lock where reads dominate overwhelmingly. The lock is the wrong tool when reads outnumber writes by orders of magnitude. Its bookkeeping costs more than a plain lock on short sections, and the starvation risk is highest precisely in the read-heavy case it is supposed to serve.

When you see it

  • A writer thread parked in `WriteLock.lock` across many consecutive thread dumps while readers churn
  • Read lock count consistently above zero, never momentarily zero, in a monitored gauge
  • Configuration, cache invalidation, or refresh tasks that apply promptly at low traffic and never at peak
  • Stale data served for minutes or hours with no error, alert, or log line
  • The writer acquires instantly the moment traffic drops, which makes the problem look like a load issue rather than a lock issue
  • For the upgrade variant: a single thread hangs on its first call, deterministically, with no deadlock reported
  • Switching the lock to fair mode fixes the starvation and visibly reduces read throughput

How to diagnose it

Step 1

Expose the lock’s own instrumentation

`ReentrantReadWriteLock` reports its state, which turns a suspicion into numbers. A read count that never touches zero with a queued writer is starvation, stated precisely.

Metrics.gauge("rw.readers", rw, ReentrantReadWriteLock::getReadLockCount);
Metrics.gauge("rw.queuedWriters", rw, ReentrantReadWriteLock::getQueuedWriterThreads);
log.info("readers={} writeLocked={} queuedWriters={}",
    rw.getReadLockCount(), rw.isWriteLocked(), rw.getQueueLength());

Step 2

Confirm the writer is queued across dumps

Successive dumps showing the same writer parked in `WriteLock.lock` while reader stacks change is the fingerprint. If the writer’s stack is identical and the readers rotate, the writer is starving rather than the whole lock being stuck.

for i in 1 2 3; do jcmd <pid> Thread.print | grep -A 5 "config-refresher"; sleep 5; done

Step 3

Measure how long the writer waits

Timestamp the acquisition attempt and the success. Any wait beyond a second or two on a background writer is starvation in progress, and having the number makes the case without argument.

long t0 = System.nanoTime();
rw.writeLock().lock();
writerWait.record(System.nanoTime() - t0, TimeUnit.NANOSECONDS);

Step 4

Distinguish starvation from a leaked read lock

Watch the read count over time. Fluctuating but never zero is starvation. Pinned at a constant value while traffic varies means a read hold was never released — a different bug with a different fix.

Step 5

Test the upgrade path explicitly

A single-threaded test that takes the read lock and then attempts the write lock, with a timeout, catches the upgrade bug deterministically. It needs no concurrency at all, which makes it cheap to keep in CI.

rw.readLock().lock();
assertFalse(rw.writeLock().tryLock(1, TimeUnit.SECONDS), "upgrade must not succeed");

Step 6

Reproduce with saturating readers

Drive continuous overlapping reads from more threads than cores and attempt a write. If the write never completes within a generous deadline, you have reproduced production behaviour in a test.

The fix

Construct the lock in fair mode when a writer must make progress: `new ReentrantReadWriteLock(true)`. Fair mode queues arriving readers behind a waiting writer, so the read count drains and the writer acquires in bounded time. Read throughput drops measurably because barging is no longer allowed — that is the price, and for a lock protecting state that must be updatable it is usually the right price.

Better in most cases: remove the writer’s need for exclusivity. Hold the state as an immutable snapshot behind a `volatile` reference — readers take no lock at all and the writer publishes a new snapshot with a single assignment. There is no read count, no queue, and no starvation to reason about. This is the fix to prefer whenever the state can be copied cheaply enough, and for configuration and routing tables it almost always can.

Never attempt to upgrade. Release the read lock, acquire the write lock, then **re-validate the condition you read** before acting on it, because the state can have changed in the gap. Skipping the re-validation swaps a hang for a lost update, which is the more common mistake once people learn upgrade is impossible.

Use `StampedLock` where reads are short and you want optimistic reading: `tryOptimisticRead` plus `validate` lets readers proceed with no lock acquisition at all in the common case, and `tryConvertToWriteLock` supports the conversion that `ReentrantReadWriteLock` cannot. It is not reentrant and does not support conditions, so it is a targeted tool rather than a drop-in replacement — using it reentrantly will self-deadlock.

Shorten read critical sections aggressively. Copy what you need out under the read lock and process outside it. Shorter reads make zero-reader windows common, which reduces starvation pressure even in non-fair mode and helps every other contention property of the lock.

If reads dominate by orders of magnitude and sections are tiny, drop the read-write lock for a `ConcurrentHashMap` or a plain lock. A read-write lock has more bookkeeping than a simple monitor, so for very short sections it can be slower *and* expose you to starvation.

As a diagnostic backstop, have the writer use `tryLock(timeout)` and log or alert when it fails. Starvation that produces a metric is a bug you find in an afternoon; starvation that produces silence is stale data nobody notices for a week.

// Starves the writer: non-fair default lets arriving readers barge past it
private final ReentrantReadWriteLock rw = new ReentrantReadWriteLock();

Config get(String k) {
    rw.readLock().lock();
    try { return serialise(config.get(k)); }   // long read section
    finally { rw.readLock().unlock(); }
}

// Self-deadlocks on the first call: a read hold can never be upgraded
rw.readLock().lock();
try {
    if (config.isStale()) {
        rw.writeLock().lock();                 // waits for itself forever
    }
} finally { rw.readLock().unlock(); }

// Fix 1: fair mode, and a short read section
private final ReentrantReadWriteLock rw = new ReentrantReadWriteLock(true);

Config get(String k) {
    Config c;
    rw.readLock().lock();
    try { c = config.get(k); } finally { rw.readLock().unlock(); }
    return serialise(c);                       // work outside the lock
}

// Fix 2 (preferred): no lock on the read path at all
private volatile Map<String, Config> config = Map.of();

Config get(String k)            { return config.get(k); }
void publish(Map<String, Config> next) { config = Map.copyOf(next); }

// Correct "upgrade": release, acquire, RE-VALIDATE
rw.readLock().lock();
boolean stale;
try { stale = config.isStale(); } finally { rw.readLock().unlock(); }
if (stale) {
    rw.writeLock().lock();
    try {
        if (config.isStale()) reload();        // re-check: state may have changed
    } finally { rw.writeLock().unlock(); }
}

// Legal direction: downgrade write -> read without releasing exclusivity
rw.writeLock().lock();
try {
    reload();
    rw.readLock().lock();                      // acquire read while holding write
} finally { rw.writeLock().unlock(); }         // now read-held only
try { use(config); } finally { rw.readLock().unlock(); }

How to stop it coming back

  • Choose fairness deliberately at construction and record why in a comment beside the field; the default is a throughput choice with a liveness cost
  • Prefer immutable snapshots behind a volatile reference for read-mostly state, so there is no writer to starve
  • Instrument writer wait time and alarm on it — starvation is silent by nature and needs a metric to be visible
  • Ban read-to-write upgrade with a review rule and a deterministic unit test using `tryLock`
  • Keep read sections to copying data out, never processing it, so zero-reader windows occur naturally
  • Never take a read lock reentrantly across a call boundary where the inner method might need to write
  • Give background writers a `tryLock` timeout so a starved refresh becomes an alert rather than stale data

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Why can a writer starve on a lock that is working correctly?

Because a write lock needs the read count to be zero, and a non-fair lock lets a new reader acquire while the lock is read-held even with a writer queued. If reads overlap continuously the count never reaches zero. The lock is behaving as documented — non-fair mode makes no ordering guarantee — and the starvation is a consequence of that policy.

Does fair mode fix it, and what does it cost?

Yes. Fair mode makes arriving readers queue behind a waiting writer, so the read count drains and the writer acquires in bounded time. The cost is real: barging is forbidden, so almost every acquisition becomes a park-and-wake handoff and read throughput falls measurably. For a lock whose writer must make progress, that is usually worth paying.

Why does upgrading from read to write deadlock?

The write lock waits for every reader to release, and the thread attempting the upgrade is itself one of those readers. It is waiting for itself. No cycle exists between two threads, so the JVM reports no deadlock — the thread simply parks forever on its first execution of that path.

Is downgrading from write to read allowed?

Yes, and it is a documented, useful pattern: acquire the write lock, mutate, acquire the read lock while still holding the write lock, then release the write lock. You end up read-holding without any window in which another writer could intervene, which is exactly what you want when you must read back what you just wrote.

Should I use `StampedLock` instead?

For short read sections where you can tolerate an optimistic read that may need retrying, yes — `tryOptimisticRead` avoids acquiring anything in the common case, and `tryConvertToWriteLock` gives you the conversion `ReentrantReadWriteLock` lacks. But it is not reentrant, does not support conditions, and its optimistic reads must be validated before you use the data. Reach for it deliberately, not as a drop-in upgrade.

How do I tell starvation from a read lock that was never released?

Watch the read count. Starvation shows it fluctuating with traffic but never hitting zero. A leaked read hold shows it pinned at a constant floor even when traffic drops — and in that case the writer will not acquire even on an idle system, which is the quickest test.

Related

Other errors engineers hit next to this one

Full error and symptom index →