Concurrency
Race conditions: interview questions and how to answer them
A race condition exists when the correctness of a result depends on the relative timing of operations that the system is free to interleave however it likes.
Written and reviewed by Sahil Srivastav
What it actually is
A race is not "two threads touching the same data" — that is merely sharing. A race is a sequence of operations that is correct when run alone and incorrect under some legal interleaving. The two canonical shapes are check-then-act, where a condition is verified and then acted on after it may have stopped holding, and read-modify-write, where a value is read, changed, and written back while another writer does the same.
The defining property is that the bug lives in the *gap*. `if (!map.containsKey(k)) map.put(k, v)` has a gap between the check and the put. `count = count + 1` has a gap between the load and the store. Nothing inside either line is wrong; what is wrong is assuming nothing happens in between, when the scheduler, the compiler, and the CPU are all entitled to put something there.
Races also occur without threads. Two processes both `stat` a file and then write it. Two HTTP requests both read a row and then update it. Two Kubernetes replicas both see no leader and both elect themselves. Anywhere two actors observe shared state and then act on that observation, the gap exists, and concurrency primitives inside one process do nothing about it.
Why it matters in production
Because the failure rate is load-dependent, which makes races the bugs that survive testing and appear in production. A window of a few microseconds is essentially never hit by a single-threaded test suite and hit thousands of times a day at a few hundred requests per second. The result is the ticket class every backend engineer recognises: "cannot reproduce locally, happens twice a week in prod, no stack trace".
Because the damage is usually silent and durable. A lost increment is invisible. A double-inserted ledger row is a reconciliation problem discovered weeks later by finance. An oversold seat becomes a customer-facing refusal at the gate. Unlike a crash, a race writes wrong data and carries on, so the cost is remediation of the data, not just a deploy.
And because the fix is almost never "add a lock" — it is deciding where the atomic boundary belongs. Interviewers use races to test exactly that judgement, because a candidate who reaches for `synchronized` on the method will happily wrap a database read-modify-write in a lock that protects one JVM out of six.
How it works
Check-then-act: the condition expires before you use it
You verify stock is available, then decrement. Between the two, another request decrements. The check was true when you made it and false when you relied on it. The fix is to make the check part of the mutating operation — a guarded UPDATE, `putIfAbsent`, `compareAndSet` — so there is no moment where the condition is known but not yet enforced.
Read-modify-write: the value you write is based on a stale read
Two threads load 5, both compute 6, both store 6. One increment vanished. Wrapping the write in a lock fixes it only if every reader-writer takes the same lock; expressing it as an atomic operation — `fetch_add`, `AtomicLong.incrementAndGet`, `SET n = n + 1` in SQL — fixes it without a lock because the read and write become indivisible.
The atomic boundary has to enclose the invariant
A lock that covers each individual field but not the relationship between them protects nothing. If the invariant is "total equals the sum of the lines", then adding a line and updating the total must be inside one critical section. Fine-grained locking that splits an invariant is a race with better-looking code.
Distributed races need the state store to arbitrate
When the racing actors are separate processes, the only thing they share is the database, the cache, or the broker. So the atomicity must live there: a unique constraint, a conditional update, a Lua script in Redis, a compare-and-set on a version. Process-local mutexes are worse than nothing, because they make the bug rarer and therefore harder to find.
Reproducing one deliberately
Do not hope for a timing coincidence — force it. A `CountDownLatch` that releases N threads simultaneously, a barrier in the middle of the gap, or a debug hook that sleeps between the check and the act will fail a racy implementation on the first run. This is also how you write a regression test that would have caught it.
Implementing it
Name the invariant before choosing a primitive. "Stock never goes negative", "one ledger row per payment", "at most one leader" — each of those points at a different mechanism, and the invariant is what the test should assert.
Prefer pushing atomicity into the store that owns the data: a unique index, a guarded UPDATE, or a CAS. It is the only place all actors can see, and it survives extra replicas, retries, and the batch job nobody told you about.
When you must lock, keep the critical section free of I/O and free of calls into code you do not control, because a lock held across a network call multiplies latency by contention.
Write the concurrency test with a latch and run it a few thousand iterations in CI. Races are probabilistic, so a single run proves nothing; a loop with a released-simultaneously barrier turns a rare bug into a deterministic failure.
// Racy: the gap between containsKey and put is legal to interleave
if (!cache.containsKey(key)) {
cache.put(key, expensiveLoad(key)); // two threads both load
}
// Atomic: the check and the insert are one operation
cache.computeIfAbsent(key, this::expensiveLoad);
// Racy across processes — the JVM lock protects one replica of six
synchronized (this) {
int qty = repo.findQty(sku);
if (qty >= n) repo.setQty(sku, qty - n);
}
// Atomic in the one place every replica shares
// UPDATE inventory SET qty = qty - :n WHERE sku = :sku AND qty >= :n
int rows = repo.decrementIfAvailable(sku, n);
if (rows == 0) throw new OutOfStockException(sku);Interview questions and how to answer them
Here is a method that checks stock and then decrements it. What is wrong, and how do you fix it?
It is check-then-act: the stock level is true at the check and possibly false by the decrement, so two concurrent callers can both pass the check and drive stock negative. The fix depends on where the state lives. If it is a database row, express it as one guarded statement — `UPDATE inventory SET qty = qty - :n WHERE sku = :sku AND qty >= :n` — and treat zero affected rows as out of stock. If it is in-process state, use an atomic compare-and-set loop or hold a lock across both steps. What I would not do is synchronise the method, because in a multi-instance deployment that protects nothing.
Why do race conditions so rarely show up in tests?
Because the failure requires a specific interleaving and the window is usually microseconds wide. A single-threaded test never produces it; a multi-threaded test run once has a low probability of hitting it. To make it deterministic you have to remove the luck — release threads from a latch so they arrive together, or inject a delay inside the gap so the interleaving is forced — and then loop it. That is how I write the regression test, and it is also how I confirm a fix rather than hoping.
Your service runs on six pods. Does a `ReentrantLock` help?
Not for shared state outside the process. Each pod has its own lock object, so six writers still race — you have reduced the concurrency factor from sixty threads to six actors. Worse, the bug now occurs rarely enough to look fixed while still corrupting data. Cross-process races need arbitration at the shared store: a unique constraint, a conditional update, a CAS on a version, or an explicit distributed lease if the work genuinely must be exclusive.
Is `volatile` enough for a counter shared by threads?
No. `volatile` makes writes visible to other threads and prevents reordering around the access, but `n++` is a load, an add, and a store — three steps with gaps. Two threads can both load the same value and both store the same result. For a counter you want `AtomicInteger.incrementAndGet`, which compiles to a lock-free compare-and-swap loop, or `LongAdder` if the counter is hot enough that CAS contention matters.
How would you find a suspected race in a running service?
Start from the corrupted data and work backwards to the invariant it violates, because that names the critical section that is missing. Then look for check-then-act or read-modify-write on that state, including paths you did not write — a batch job, an admin endpoint, a retry. Thread dumps help for deadlock but rarely for races; what helps is a database-side check (a constraint that would have rejected the bad row) and a forced-interleaving test that reproduces it locally. Adding the missing constraint is often worth doing on its own, because it converts silent corruption into a loud error.
Answers that lose the round
- Defining a race as "two threads accessing shared data", which describes every concurrent program rather than the bug
- Reaching for `synchronized` on a method when the shared state is a database row visible to every instance of the service
- Making the window smaller — moving the check closer to the act — and calling it fixed; a smaller window is a rarer bug, not a correct program
- Adding `volatile` to fix a read-modify-write: it guarantees visibility, not atomicity, so `volatile int n; n++` is still racy
- Locking two different objects on two paths that touch the same state, so mutual exclusion never actually holds
- Claiming the test suite proves absence of races when the test is single-threaded or runs the concurrent case once
- Using a `ConcurrentHashMap` and then composing two of its operations, which is thread-safe per call and racy as a sequence
Practise race conditions in a real repository
Gronex ships this as a runnable repository: an inventory endpoint that oversells because the availability check and the decrement are separate steps. The test suite releases concurrent purchases from a latch, so the race is deterministic and a narrower window does not pass.
FAQ
Is a data race the same as a race condition?
No, and the distinction matters in interviews. A data race is the narrow memory-model notion: unsynchronised concurrent access to the same location where at least one access is a write, which is undefined behaviour in C++ and produces no visibility guarantee in Java. A race condition is the broader correctness notion: the outcome depends on timing. You can have a race condition with no data race — two perfectly synchronised database calls in the wrong order — and that is the kind most backend bugs are.
Do immutable objects eliminate races?
They eliminate races on the object's own contents, which removes a large class of bugs. They do not eliminate races on the *reference*: swapping which immutable snapshot is current is still a write, and a compare-and-set or an atomic reference is what makes that swap safe.
Can a single-threaded event loop have races?
Yes. In Node.js or an async Python service, every `await` is a yield point, so state observed before an await may have changed by the time execution resumes. There is no data race because there is one thread, but check-then-act across an await is exactly the same bug with the same fix.