Concurrency

Missed signal / lost wakeup with wait() and notify()

Written and reviewed by Sahil Srivastav

Concurrencywait/notifyLiveness
"queue-consumer-2" #31 prio=5 os_prio=0 tid=0x00007f3a180c4000 nid=0x4e2b in Object.wait() [0x00007f39ef7fd000]
   java.lang.Thread.State: WAITING (on object monitor)
	at java.lang.Object.wait(java.base@21/Native Method)
	- waiting on <0x00000006c2a41f30> (a com.example.queue.TaskQueue)
	at java.lang.Object.wait(java.base@21/Object.java:366)
	at com.example.queue.TaskQueue.take(TaskQueue.java:41)
	- locked <0x00000006c2a41f30> (a com.example.queue.TaskQueue)
	at com.example.queue.Consumer.run(Consumer.java:19)

# meanwhile:
[monitor] queue.size()=847 consumers.waiting=4 tasks.completed=0 (last 60s)

What this error actually means

A monitor signal is not a message and it is not queued. `notify()` wakes a thread that is *already waiting at this instant*; if nobody is waiting, the call does nothing at all and leaves no trace. The signal is not buffered for a future waiter, so a producer that signals a microsecond before the consumer reaches `wait()` has signalled into the void — and the consumer then sleeps on a queue it could have drained immediately.

This is why the condition and the wait must be evaluated under the same lock, and why the wait must sit inside a loop. Holding the monitor across the check-and-wait closes the race window: `wait()` atomically releases the monitor and enqueues the thread, so no producer can slip a `notify()` between your test and your sleep. Testing the condition with an `if` instead of a `while` reopens a different hole, because between being woken and re-acquiring the monitor the thread is *not* holding it, and another consumer can take the item that the signal was about.

The loop is also mandatory for a reason independent of your code. `Object.wait()` is specified to be allowed to return without any `notify`, `notifyAll`, interrupt, or timeout — a spurious wakeup. The JVM permits it because monitor implementations are built on OS primitives whose wait operations may return early, and requiring the JVM to filter those out would cost every wait a re-validation it cannot perform, since only your code knows what the condition is. So the contract puts the re-check on the caller: treat a return from `wait()` as "something may have changed", never as "the condition now holds".

There is no exception for any of this. The failure is a thread that stops doing work while every metric that measures errors stays clean. Throughput goes to zero, queue depth climbs, and nothing is logged.

Causes, most common first

  1. 1Condition tested with `if` instead of `while`. The most common form by a wide margin. After being woken, a thread must re-acquire the monitor, and during that gap another consumer can consume the item. The woken thread then proceeds as though the condition holds — returning null, decrementing below zero, or reading a slot that is now empty.
  2. 2Signal sent before the waiter arrives. Producer and consumer start concurrently; the producer enqueues and signals during the window before the consumer’s first `wait()`. Nothing is buffered, so the consumer sleeps on non-empty state. Classic at start-up and after any consumer restart, which is why it often looks like a "first run" bug.
  3. 3State mutated outside the monitor that guards the wait. The queue is updated without holding the same lock the waiter waits on — or the field is updated under lock A while the wait is on lock B. The signal and the state change are then not atomic with respect to the waiter, and no amount of `notifyAll` repairs it.
  4. 4`notify()` where `notifyAll()` is required. `notify()` wakes one arbitrary waiter. If several threads wait on the same monitor for *different* conditions — producers waiting for space and consumers waiting for items — the runtime can wake a thread whose condition is still false. That thread loops and sleeps again, and the signal that should have gone to the other kind of waiter is consumed and gone.
  5. 5Signalling outside the lock, or after releasing it. Moving `notify()` out of the synchronized block to "reduce contention" throws `IllegalMonitorStateException` if the monitor is not held, and where the code is restructured so it is legal, it reintroduces the gap between the state change and the signal.
  6. 6Interrupt swallowed inside the wait loop. `wait()` throws `InterruptedException` and clears the interrupt flag. Catching and continuing turns a shutdown request into an infinite loop; catching and ignoring it makes the thread unstoppable, so the pool never terminates and the stall looks like a lost wakeup.

When you see it

  • Queue depth grows without bound while consumer threads sit in `WAITING (on object monitor)`
  • A single unrelated item arriving later wakes one consumer, which drains a burst and then sleeps again — progress in stutters
  • CPU is near idle during the stall; nothing is spinning, nothing is erroring
  • It reproduces under load and essentially never in a single-producer, single-consumer test
  • One consumer of several is permanently asleep while the others work, because `notify()` woke the wrong one
  • A thread dump taken twice shows byte-identical frames, yet the JVM reports no deadlock — the threads are waiting, not blocked in a cycle

How to diagnose it

Step 1

Separate waiting from blocked in the dump

This distinction is the whole diagnosis. `WAITING (on object monitor)` with an `Object.wait` frame means the thread released the lock and is waiting for a signal — a lost wakeup or an unmet condition. `BLOCKED (on object monitor)` means it is waiting to acquire, which is contention or a leaked lock instead.

jcmd <pid> Thread.print | grep -E "Thread.State|Object.wait|- waiting on|- locked"

Step 2

Compare queue depth with waiter count at the same moment

The decisive observation. Non-empty state plus waiting consumers plus no throughput is a lost wakeup, definitionally. Expose both as gauges so you can read them together rather than inferring from logs.

Metrics.gauge("queue.depth", q, TaskQueue::size);
Metrics.gauge("queue.waiters", q, TaskQueue::waiterCount);

Step 3

Confirm the state is guarded by the monitor you wait on

Read the dump line `- waiting on <0x...> (a com.example.queue.TaskQueue)` and check that every mutation of the queue happens inside `synchronized` on that same object. A mismatch here is the bug regardless of how the loop is written.

Step 4

Reproduce with a deliberately adversarial schedule

Start the producer first, let it enqueue and signal, and only then start the consumer. A single-item, ordered reproduction is far more reliable than load, because load hides the window you are trying to hit.

q.put(task);                 // signal with no waiter present
Thread.sleep(50);
new Thread(consumer).start(); // must still find the task

Step 5

Force spurious wakeups to test the loop

Call `notifyAll()` periodically from a background thread while the condition is false. Correct code loops and sleeps again; code with an `if` will fall through and misbehave, which is exactly the latent bug you want surfaced in CI.

The fix

Wait inside a `while` that re-tests the condition, always, with no exceptions. The shape is: acquire the monitor, `while (!condition) wait();`, then act. `if (!condition) wait();` is wrong even when the logic looks airtight, because the specification permits a wakeup you did not send and because re-acquisition is not atomic with being woken.

Publish the state change and the signal under the same monitor that the waiter waits on. Mutate, then signal, then leave the block. This makes the transition atomic from the waiter’s point of view and closes the pre-arrival race, because a producer holding the monitor cannot signal while a consumer is between its check and its `wait()`.

Use `notifyAll()` unless you can prove every waiter on that monitor waits for the identical condition and any one of them can consume the event. The cost of `notifyAll` is some redundant wakeups that immediately loop and sleep again; the cost of a wrong `notify` is a permanent stall. That trade is not close.

Better: stop hand-rolling it. `ArrayBlockingQueue` or `LinkedBlockingQueue` give you `put`/`take` with the condition loops, signalling, and interrupt handling already correct, and a bounded queue additionally gives you backpressure. Replacing a bespoke monitor queue with a `BlockingQueue` removes the entire bug class rather than fixing this instance of it.

If you need multiple distinct conditions on one lock, move to `ReentrantLock` with a separate `Condition` per predicate — `notFull` and `notEmpty` — and `signalAll` the specific one. Distinct conditions are precisely what a single monitor cannot express, and they are why the wrong-waiter variant exists at all.

Add a timed wait as a liveness net once the logic is right: `wait(timeoutMillis)` in the loop bounds a lost signal to one timeout instead of forever, and a counter of timeout-driven iterations gives you a metric that a signal was missed. Do not use it as the fix — a loop that only makes progress on timeouts is a broken design with a heartbeat.

// Broken three ways: `if` instead of `while`, and notify() to mixed waiters
synchronized Task take() throws InterruptedException {
    if (items.isEmpty()) wait();      // spurious or stolen -> returns on empty
    return items.removeFirst();       // NoSuchElementException, or worse, silence
}
synchronized void put(Task t) {
    items.addLast(t);
    notify();                          // may wake a producer waiting for space
}

// Correct: loop on the predicate, mutate and signal under the same monitor
synchronized Task take() throws InterruptedException {
    while (items.isEmpty()) wait();
    Task t = items.removeFirst();
    notifyAll();                       // space is now available
    return t;
}
synchronized void put(Task t) throws InterruptedException {
    while (items.size() == capacity) wait();
    items.addLast(t);
    notifyAll();
}

// Better: the library already has this correct, with backpressure included
BlockingQueue<Task> q = new ArrayBlockingQueue<>(1_000);
q.put(task);                            // blocks when full
Task t = q.take();                      // blocks when empty

How to stop it coming back

  • Treat `if (...) wait();` as a defect on sight — the condition loop is not a style preference, it is the contract
  • Prefer `java.util.concurrent` queues and synchronisers over hand-written monitor protocols in all application code
  • Inject spurious `notifyAll()` calls in tests so any code that skips the re-check fails in CI rather than in production
  • Export queue depth and waiter count as separate gauges; their combination is the only cheap detector for this failure
  • Alarm on "depth above zero and throughput zero" rather than on error rate, because this failure produces no errors
  • Never mutate guarded state outside the monitor the waiters wait on, and keep the state and its lock in the same class so the pairing is enforceable

Practise this failure in a real repository

Gronex ships this as a runnable repository: a bespoke queue whose consumers test the predicate once and whose producer signals a single arbitrary waiter. The test suite starts producers before consumers, injects spurious wakeups, and asserts every enqueued task is eventually processed — so only a correct condition loop and signalling discipline passes.

FAQ

Why are spurious wakeups allowed at all?

Monitor waiting is implemented on operating-system primitives whose wait operations may return early, and the JVM cannot re-validate your condition on your behalf — only your code knows what it is. Rather than paying to filter early returns on every wait, the specification places the re-check on the caller. That is the entire reason `wait()` must be called in a loop.

Where does the signal go if nobody is waiting?

Nowhere. `notify()` and `notifyAll()` have no effect and no memory when the wait set is empty; there is no counter and no buffer. This is the core difference from a `Semaphore`, whose `release()` increments a permit count that a later `acquire()` will consume — which is why a semaphore is the right tool when the event must survive the absence of a waiter.

Is `notifyAll()` not a performance problem?

It wakes every waiter, each of which re-acquires the monitor, re-tests its predicate, and usually sleeps again — a "thundering herd" that costs context switches. It is a real cost and almost always the right trade against a permanent stall. When the herd genuinely matters, the fix is per-condition `Condition` objects so you can signal precisely, not reverting to `notify()`.

Does `ReentrantLock` with `Condition` remove the need for the loop?

No. `Condition.await()` is documented to be subject to spurious wakeup exactly as `Object.wait()` is, and the same stolen-item race exists between being signalled and re-acquiring the lock. The loop is required there too. What `Condition` buys you is multiple independent wait sets on one lock.

How do I tell a lost wakeup from a deadlock in a dump?

Read the thread state. A lost wakeup shows `WAITING (on object monitor)` with an `Object.wait` frame and a `- locked` line for the same monitor, because `wait()` releases it. A deadlock shows `BLOCKED (on object monitor)` threads and, usually, a `Found one Java-level deadlock` section naming the cycle. Waiting threads hold nothing; blocked threads are trying to acquire.

Related

Other errors engineers hit next to this one

Full error and symptom index →