Concurrency

Visibility bug: a non-volatile stop flag

Written and reviewed by Sahil Srivastav

ConcurrencyJava memory modelVisibility
# stop() was called at 14:02:11. The worker is still running 40 minutes later.
"ingest-worker" #18 prio=5 os_prio=0 tid=0x00007f4b1c0a9800 nid=0x2f11 runnable [0x00007f4ae8bfe000]
   java.lang.Thread.State: RUNNABLE
	at com.example.ingest.Worker.run(Worker.java:27)
	at java.base/java.lang.Thread.run(Thread.java:1583)

# The loop reads a non-volatile boolean. After C2 compiled it, the read was
# hoisted out of the loop:
$ java -XX:+UnlockDiagnosticVMOptions -XX:+PrintCompilation Worker
    142  418 % 4  com.example.ingest.Worker::run @ 6 (41 bytes)
# ... and the process never exits, because the worker thread is non-daemon.

What this error actually means

One thread wrote `running = false`. Another thread reads `running` in a loop and never observes the write. There is no exception, no error, and no way to tell from the source that anything is wrong — the write happened and the read is right there.

The Java memory model does not promise that a write by one thread becomes visible to another unless there is a *happens-before* relationship between them. A plain write to a plain field creates no such relationship. Without one, the compiler, the JIT, and the hardware are all free to keep the value where it is cheapest: a register, a store buffer, a core-local cache line. Nothing is obliged to make it globally visible, ever.

The specific mechanism that usually turns "eventually visible" into "never visible" is the JIT, not the CPU. When C2 compiles a loop whose condition reads a non-volatile field that the loop body never modifies, it can treat that read as loop-invariant and hoist it out — loading the value once before the loop and then testing a register. That transformation is legal precisely because the absence of a happens-before edge means the compiler is entitled to assume no other thread is changing it. The result is `while (true)` in the compiled code, and no later write by any thread can affect it.

This also explains the most misleading property of the bug: it works until it does not. Interpreted and tier-1 compiled code re-reads the field each iteration, so the flag appears to work during start-up and in short tests. Once the method is hot enough for C2 — typically after several thousand iterations — the optimised version takes over and the loop becomes uninterruptible. A load test that runs for ten seconds passes; the same code in production never stops.

Declaring the field `volatile` fixes it because a volatile write and a subsequent volatile read of the same field establish a happens-before edge. The compiler may no longer cache the value across the loop, and the hardware is instructed to make the write visible. That is the entire fix — and it is a correctness fix, not a hint.

Causes, most common first

  1. 1The flag is a plain `boolean` field. The bug in one line. A non-volatile field read in a loop condition and written by another thread has no ordering or visibility guarantee, and is a legal candidate for hoisting out of the loop.
  2. 2`volatile` on the reference but not the field being read. A volatile reference to a holder object whose `running` field is plain. The reference read is ordered; the subsequent field read is not. The guarantee applies to the specific field declared volatile, not transitively to everything reachable from it.
  3. 3The flag is read once into a local before the loop. A manual version of the same hoist, often written deliberately as an optimisation. Here the source is unambiguous, which at least makes it findable by reading.
  4. 4Written under a lock and read without one. The write is inside a `synchronized` block, so the author believes it is published. Publication requires both sides to synchronise on the same monitor: a plain read outside the lock establishes no happens-before edge with the write inside it.
  5. 5A long compute loop with no blocking call. Related and frequently co-occurring: a loop with no blocking operation also receives no `InterruptedException`, so interruption cannot stop it either. Both the flag and the interrupt are invisible, and the task is unstoppable by any means.
  6. 6A composite condition where only part is volatile. `while (running && !paused)` with one field volatile and the other not. The volatile read forces a re-read of that field and prevents hoisting of the loop as a whole, which can make the non-volatile field *appear* to work — until an unrelated change to the loop removes the volatile access and the latent bug surfaces.

When you see it

  • A worker thread continues after its stop method was called, with no error logged
  • The JVM will not exit, because the surviving worker thread is non-daemon
  • The process ignores SIGTERM and is SIGKILLed by the orchestrator after the grace period
  • Adding a log line, a `sleep`, or a `synchronized` block inside the loop makes it start working — which is a symptom, not a fix
  • It stops correctly in debug or with `-Xint` and fails when running normally, because the JIT is what introduces the hoist
  • Reproduces readily on AArch64 and less often on x86, since a strong hardware memory model hides part of the problem
  • CPU shows one core pinned by a loop that should have ended

How to diagnose it

Step 1

Read the field declaration first

For a thread that will not stop, check the modifier on the flag before anything else. A plain `boolean` written by one thread and read by another in a loop is the answer, and confirming it takes seconds.

grep -rn -E "(private|protected|public)? *(static)? *boolean +(running|stopped|shutdown|active)" src/main/java | grep -v volatile

Step 2

Confirm the loop is running, not blocked

A `RUNNABLE` thread sitting in your own loop frame across multiple dumps means it is spinning on a condition that will never change. A `WAITING` or `BLOCKED` thread is a different problem entirely — a lost wakeup or contention.

jcmd <pid> Thread.print | grep -A 4 "ingest-worker"

Step 3

Establish that the JIT is the trigger

Run the reproducer interpreted. If it terminates with `-Xint` and hangs without it, the hoist is confirmed and you do not need to reason about hardware memory models at all.

java -Xint Worker      # terminates
java Worker            # hangs once C2 compiles the loop

Step 4

Inspect the compiled code if you need proof

With hsdis available, the compiled loop shows the field load outside the loop body and a register test inside it. This is the most direct evidence there is, and useful when convincing someone that the source "obviously works".

java -XX:+UnlockDiagnosticVMOptions -XX:+PrintAssembly \
     -XX:CompileCommand=print,com.example.ingest.Worker::run Worker | head -80

Step 5

Test termination with a deadline

The regression test that actually matters: run the worker long enough to be compiled, set the flag, and assert the thread dies within a deadline. Without the warm-up iterations the test passes even with the bug present.

worker.start();
Thread.sleep(2_000);            // let C2 compile the loop
worker.stop();
assertTrue(worker.joinFor(Duration.ofSeconds(2)), "worker ignored stop flag");

The fix

Declare the flag `volatile`. A volatile write followed by a volatile read of the same field creates the happens-before edge the JMM requires, and it forbids the compiler from caching the value across loop iterations. For a single boolean flag with no other invariants attached, this is the complete and correct fix — not a mitigation.

Prefer interruption to a custom flag for stopping threads. `Thread.interrupt()` plus a loop condition of `!Thread.currentThread().isInterrupted()` uses the platform mechanism, which also unblocks interruptible blocking calls. A hand-rolled flag cannot wake a thread that is parked in `queue.take()`; interruption can. Use both together: the flag for a graceful drain, interruption for the forced stop.

Use `AtomicBoolean` when you need to *act once* on the transition, not merely observe it. `compareAndSet(false, true)` guarantees exactly one caller performs the shutdown work, which a volatile boolean does not — and it carries the same visibility guarantees. Do not reach for it purely for visibility, where volatile is cheaper and clearer.

If the flag is read under a lock, write it under the *same* lock. Publication through a monitor works only when both sides use it; mixing a synchronised write with a plain read is the most common way to get this wrong while believing it is handled.

Never read the flag once into a local outside the loop, and treat any composite condition with mixed volatility as broken. Make every field involved in the loop condition volatile, so the correctness of the loop does not depend on which field happens to force a re-read.

Do not rely on `Thread.sleep`, logging, or a `synchronized` block inside the loop to make it work. These often do make the symptom disappear — a method call can inhibit the hoist and a synchronised block inserts a barrier — but they provide no guarantee, and the next refactor that removes the log line reintroduces a bug nobody will connect to it. Fix the declaration.

Make long-lived worker threads daemon threads *in addition* to fixing the flag, so a future stop bug degrades into a lost background task at exit rather than a process that cannot terminate. This is a backstop for exit, never a substitute for a working stop path.

// The loop may be compiled to `while (true)`: no happens-before edge exists,
// so the read is loop-invariant as far as the compiler is concerned.
private boolean running = true;

public void run() {
    while (running) {          // hoisted: value read once, then a register test
        ingestBatch();
    }
}
public void stop() { running = false; }   // write may never become visible

// Also broken: published under a lock, read without one
synchronized void stop() { running = false; }
public void run() { while (running) ingestBatch(); }   // plain read, no edge

// Correct: volatile establishes the happens-before edge and forbids hoisting
private volatile boolean running = true;

public void run() {
    while (running && !Thread.currentThread().isInterrupted()) {
        ingestBatch();
    }
}
public void stop() { running = false; }

// Better for a stoppable worker: the flag drains, the interrupt unblocks
public void shutdown() {
    running = false;      // graceful: finish the current batch and exit
    thread.interrupt();   // forced: also unblocks queue.take() and sleep()
}

// Use AtomicBoolean when exactly one caller must perform the shutdown work
private final AtomicBoolean stopped = new AtomicBoolean();
public void shutdown() {
    if (stopped.compareAndSet(false, true)) closeResources();
}

How to stop it coming back

  • Treat any field written by one thread and read by another without synchronisation as a defect; `volatile` or a lock, never neither
  • Prefer interruption over bespoke flags for thread termination, since it also unblocks blocking calls
  • Write termination tests that warm the loop for seconds before setting the flag, so the JIT-compiled version is the one under test
  • Run concurrency suites on AArch64 as well as x86 — a strong hardware memory model masks a large fraction of visibility bugs
  • Never read a loop-control field into a local before the loop, and keep loop conditions built only from volatile or lock-protected state
  • Make long-lived worker threads daemon threads as a backstop so a stop bug cannot prevent process exit
  • Alarm on pods being SIGKILLed after the termination grace period; an unstoppable loop is a common and easily missed cause

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Why does the thread never see a write that definitely happened?

Because without a happens-before edge the JMM places no obligation on when — or whether — a plain write becomes visible to another thread. The value may live in a register or a store buffer indefinitely, and more decisively, the JIT may hoist the read out of the loop entirely, so the reading thread stops consulting memory at all.

Why does adding a `Thread.sleep(100)` to the loop fix it?

It usually does, and it is not a fix. A call that the compiler cannot prove side-effect-free tends to inhibit hoisting the read, so the loop re-reads the field each iteration. That is an artefact of this version of the optimiser, guaranteed by nothing, and it will be silently undone by any future change. The same applies to a log statement or a synchronised block.

Is `volatile` enough, or do I need `AtomicBoolean`?

`volatile` is sufficient for a flag that one thread sets and others observe — it provides exactly the visibility and ordering required. `AtomicBoolean` adds atomic read-modify-write, which you need only when you must guarantee that exactly one caller performs an action on the transition, such as running cleanup once.

Does `volatile` make the field thread-safe for counting?

No. It guarantees visibility and ordering, not atomicity. `count++` on a volatile field is still a read, an increment, and a write, and two threads can interleave to lose an update. Use `AtomicInteger`, or `LongAdder` under contention.

Why does it reproduce on Apple silicon or Graviton and not on my x86 laptop?

x86 has a strong memory model in which stores become visible in order with little extra work, which hides the hardware half of the problem. AArch64 is weakly ordered and requires explicit barriers, so a missing `volatile` shows up far more readily. The JIT hoisting half happens on both — a strong CPU does not protect you from it.

Would making the field `final` or the object immutable help?

Not for a mutable flag, by definition — a stop flag must change. Immutability is the right answer for shared *state* and not for a control signal. For the signal, use `volatile` or interruption; for the state the worker processes, immutability plus a single volatile publishing write is the pattern that scales.

Related

Other errors engineers hit next to this one

Full error and symptom index →