Concurrency
Race on lazy initialisation of a shared cache
Written and reviewed by Sahil Srivastav
2026-06-18T11:02:04.118Z INFO PriceCache - loading price book for region=eu-west (cache miss)
2026-06-18T11:02:04.119Z INFO PriceCache - loading price book for region=eu-west (cache miss)
2026-06-18T11:02:04.119Z INFO PriceCache - loading price book for region=eu-west (cache miss)
... 47 identical lines within 3ms ...
2026-06-18T11:02:06.884Z WARN HikariPool-1 - Connection is not available, request timed out after 30000ms
2026-06-18T11:02:09.211Z ERROR MeterRegistry - duplicate gauge registration: pricebook.entries{region=eu-west}
# 47 threads, 47 identical queries, 47 connections, one surviving cache entry.What this error actually means
A lazily populated cache has a window between "the value is missing" and "the value is stored". Under concurrency, every thread that arrives during that window sees a miss and starts its own load. If the load takes 200 milliseconds and the endpoint serves 200 requests per second for that key, roughly forty threads load it simultaneously, and thirty-nine of the results are thrown away.
What turns a wasteful race into an outage is what the load *consumes* rather than what it returns. Forty concurrent loads take forty connections from a pool sized for twenty, so the pool times out and unrelated endpoints start failing. Or the load registers a metric, opens a client, or starts a refresh thread, and thirty-nine of those side effects are created and then abandoned — leaked, because nobody holds the reference that would let them be closed. The duplicate-gauge error in the output above is the cache stampede telling you it also leaked thirty-nine objects.
The timing is characteristically nasty. It happens at start-up, when every key is cold at once; after a deployment, when a fresh instance takes full traffic with an empty cache; and at expiry, when a popular entry becomes invalid and the herd re-forms on a warm system with no warning. That last one is why a cache that has been stable for months can cause an incident at a fixed interval matching its TTL.
The distinction from a plain check-then-act race matters for the fix. Making the insertion atomic stops duplicate *entries*; it does not, on its own, stop duplicate *work*. What you need is for the first thread to install a placeholder that later threads can wait on, so exactly one load happens and everyone gets its result.
Causes, most common first
- 1`get`-then-`load`-then-`put` around a cache. The direct form. Three separate operations with the load in the middle, so every thread arriving before the `put` repeats the whole load. Making the map concurrent changes nothing, because the window is in your code rather than in the map.
- 2Expiry with no coalescing. The entry is evicted or times out and the herd forms instantly on a busy key. This is the version that hits production rather than staging, because it needs sustained traffic on a single hot key — exactly what a warm system has and a test does not.
- 3A load with side effects that are not idempotent. The loader opens a connection, registers a gauge, subscribes to a topic, or starts a refresher. The discarded duplicates keep their side effects. This converts a performance bug into a resource leak and, for subscriptions, into duplicate message processing.
- 4`putIfAbsent` with an eagerly evaluated value. The insertion is atomic but the argument is computed before the call, so every thread still performs the full load. The cache ends up correct and the source system still receives N queries — which is why this "fix" is so often applied and so often does not help.
- 5Per-key locks taken after the load rather than before. The lock protects the write instead of the read-load-write sequence. Threads serialise on the insertion, having each already done the expensive part in parallel.
- 6Negative results not cached. When the load legitimately finds nothing, nothing is stored, so every subsequent request repeats it. A stampede on a non-existent key runs forever rather than resolving after the first load, and it is invisible in cache hit-rate metrics because there is never a hit to measure.
When you see it
- Many identical "cache miss, loading" log lines within a few milliseconds for the same key
- Connection-pool timeouts or downstream rate limiting that begin seconds after a deploy or a restart
- Load on the source system spiking at an interval that matches the cache TTL exactly
- Duplicate-registration errors from metric registries or JMX for a resource that should be unique
- A slow leak of connections, clients, or threads proportional to concurrency at the moment of each miss
- Latency percentiles with a sharp cliff: the first request is slow and so are the forty behind it, rather than one slow and the rest fast
- Total throughput falling as concurrency rises, because the herd competes with itself for the source
How to diagnose it
Step 1
Count loads against distinct keys
The decisive measurement, and it needs no instrumentation beyond a log line at the start of every load. More loads than misses of distinct keys is stampede, quantified. The ratio also tells you the concurrency at the moment of the miss.
grep "cache miss" app.log | awk '{print $NF}' | sort | uniq -c | sort -rn | headStep 2
Correlate load count with source-side load
Check the source for bursts of identical queries. In PostgreSQL, `pg_stat_statements` will show one statement with a call count far above the number of logical misses, which attributes the load spike to the cache rather than to traffic.
SELECT calls, mean_exec_time, left(query, 60) FROM pg_stat_statements ORDER BY calls DESC LIMIT 10;Step 3
Check whether the spike period matches the TTL
Plot source-side load and look for a regular sawtooth. A period equal to the cache TTL is conclusive: the herd is re-forming at each expiry, and the fix is coalescing plus refresh-ahead rather than anything about traffic.
Step 4
Reproduce with a barrier
Release N threads simultaneously against one cold key and count invocations of the loader. Correct coalescing yields exactly one. This is a fast unit test and it becomes the regression test for the fix.
var start = new CountDownLatch(1);
var calls = new AtomicInteger();
// N threads: start.await(); cache.get("k"); loader increments calls
start.countDown();
assertEquals(1, calls.get());Step 5
Look for abandoned side effects
If the loader creates anything, count live instances against distinct keys in a heap histogram. An excess is the leak, and its size tells you how much of the incident was memory rather than latency.
jcmd <pid> GC.class_histogram | grep -E "PriceBook|Subscription|Client"The fix
Install a placeholder, not a value. Store a `CompletableFuture` (or a `FutureTask`) per key so the first thread to miss publishes a promise immediately and performs the load; every later thread finds the future and waits on it. Exactly one load runs, all callers get the same result, and the window that caused the stampede no longer exists because the placeholder is published before the work starts. This is the classic memoiser pattern and it is the core of the fix.
Make the placeholder installation atomic with `computeIfAbsent`, and start the load *outside* the mapping function. The function should do nothing but create the future — it runs while the map’s bin lock is held, so performing the load inside it blocks every other operation on that bin and can throw `IllegalStateException: Recursive update` if the loader touches the same cache.
Remove the failed future on failure, or the cache will memoise the error forever. `whenComplete((v, ex) -> { if (ex != null) cache.remove(key, future); })` is the smallest correct version — and it must remove *that specific future* with the two-argument `remove`, not whatever is in the map now, or you will evict a successful reload performed by another thread.
Prefer a cache library that already does all of this. Caffeine’s `AsyncLoadingCache` coalesces concurrent loads per key, bounds the cache, handles failure eviction, and supports refresh-ahead. Hand-rolling the memoiser is instructive; shipping the hand-rolled version and then discovering the failure-eviction and bounding requirements one incident at a time is not.
Use refresh-ahead rather than expire-and-reload for hot keys. `refreshAfterWrite` reloads in the background while continuing to serve the stale value, so an expiry never exposes a window in which the key is missing. This eliminates the TTL-periodic stampede specifically, which is the form that hits production.
Do not synchronise the whole cache as the fix. A single lock around the accessor does stop the duplicate loads, and it also serialises every hit on every key behind whichever thread is currently performing a slow load. You trade a burst of duplicate work for a global choke point, which under load is usually worse than the problem.
Cache negative results with a short TTL so a missing key does not stampede indefinitely, and make loaders idempotent and side-effect-free. If a load must create a resource, the coalescing above guarantees it happens once — which is the real reason to fix this properly rather than tolerate the waste.
// Stampede: the window between the miss and the put is the whole bug
PriceBook pb = cache.get(region);
if (pb == null) {
pb = loadPriceBook(region); // 47 threads all get here
cache.put(region, pb);
}
// Atomic insert, still 47 loads: the argument is evaluated eagerly
cache.putIfAbsent(region, loadPriceBook(region));
// Coalesced: publish a promise first, then do the work exactly once
private final ConcurrentHashMap<String, CompletableFuture<PriceBook>> cache =
new ConcurrentHashMap<>();
PriceBook get(String region) {
CompletableFuture<PriceBook> f = cache.computeIfAbsent(region, r -> {
var promise = new CompletableFuture<PriceBook>();
loaderPool.execute(() -> { // load OUTSIDE the bin lock
try { promise.complete(loadPriceBook(r)); }
catch (Throwable t) { promise.completeExceptionally(t); }
});
return promise;
});
// do not memoise failures: remove this exact future, not whatever is there now
f.whenComplete((v, ex) -> { if (ex != null) cache.remove(region, f); });
return f.join();
}
// Preferred in production: the library handles coalescing, bounds,
// failure eviction and refresh-ahead
AsyncLoadingCache<String, PriceBook> prices = Caffeine.newBuilder()
.maximumSize(10_000)
.refreshAfterWrite(Duration.ofMinutes(5)) // reload behind a served value
.expireAfterWrite(Duration.ofMinutes(30))
.buildAsync((region, executor) ->
CompletableFuture.supplyAsync(() -> loadPriceBook(region), executor));How to stop it coming back
- Log every load with its key and alarm when loads exceed distinct-key misses — the ratio is the stampede metric
- Use a cache library with per-key load coalescing rather than a bare map plus a loader
- Prefer `refreshAfterWrite` over bare expiry for hot keys so no window exists in which a popular key is absent
- Assert single-invocation in a test: N threads released on a barrier against one cold key must call the loader exactly once
- Keep loaders idempotent and free of registration side effects; if a load must create a resource, make coalescing a hard requirement
- Cache negative lookups briefly so a missing key cannot stampede indefinitely and invisibly
- Warm known-hot keys at start-up so the first production request after a deploy is not the one that forms the herd
FAQ
Does `computeIfAbsent` on its own prevent a cache stampede?
It prevents duplicate *entries* and, because the mapping function runs under the bin lock, it does serialise loads for the same key. The trouble is that a slow load then holds that bin lock, blocking unrelated keys in the same bin and risking `Recursive update` if the loader touches the cache. Publishing a future and loading outside the lock gets the coalescing without holding the map hostage.
Why not just synchronise the whole accessor?
Because it converts a burst of wasted work into a permanent global bottleneck. Every cache hit for every key then queues behind whichever thread is performing the slowest current load. Under the load that triggers stampedes in the first place, that is usually a worse outage than the stampede.
What happens if the load fails and I cached the future?
Every subsequent caller joins the failed future and receives the same error, permanently — a single transient failure becomes a sticky one. Remove the failed future on completion, using the two-argument `remove(key, future)` so you only evict that specific attempt and not a successful reload another thread has since installed.
Why does the stampede happen at the same time every day?
Because it is driven by expiry rather than by traffic. A hot key written at a fixed time expires a TTL later and the herd re-forms, producing a sawtooth on the source system with a period equal to the TTL. Refresh-ahead removes it by reloading before the value becomes unavailable.
Is this the same bug as double-checked locking?
They are relatives with different consequences. Double-checked locking without `volatile` is a *visibility* bug — the object is published unsafely and can be seen half-built. A stampede is an *atomicity* bug — the value is published correctly, just many times. You can have either without the other, and the fixes differ: publication ordering versus load coalescing.
Related
Other errors engineers hit next to this one
- Intermittent 502 after an idle keep-alive connection
- SSL certificate problem: unable to get local issuer certificate
- ERR_INCOMPLETE_CHUNKED_ENCODING
- Request timeouts cascading into pool exhaustion
- Connection reset by peer on a long-polling endpoint
- OOMKilled — container exit code 137
- CrashLoopBackOff
- ImagePullBackOff / ErrImagePull