Performance
Cache stampede: interview questions and practical design
A cache stampede occurs when many callers miss or expire the same key and simultaneously load the origin, overwhelming the very dependency the cache was meant to protect.
Written and reviewed by Sahil Srivastav
What it actually is
A cache stampede occurs when many callers miss or expire the same key and simultaneously load the origin, overwhelming the very dependency the cache was meant to protect.
A popular key can turn one expiry into a synchronized burst. The origin then slows, misses last longer, and more callers pile on.
The useful interview answer is precise about the boundary: Elect one caller to load a missing key while others wait for the same result. The lock or promise must have a timeout and ownership so a crashed loader does not block all readers.
Why it matters in production
A popular key can turn one expiry into a synchronized burst. The origin then slows, misses last longer, and more callers pile on.
Avoiding the stampede requires coordinating misses and expiry, not simply increasing cache capacity or adding a longer TTL.
How it works
Single-flight loading
Elect one caller to load a missing key while others wait for the same result. The lock or promise must have a timeout and ownership so a crashed loader does not block all readers.
Jittered expiry
Spread expiration times so a batch of keys does not expire together. Jitter reduces synchronisation but does not solve a single hot key by itself.
Stale-while-revalidate
Serve a still-usable stale value while one background refresh updates it. Set a hard staleness limit and ensure refresh work is bounded.
Implementing it
Expire one hot key under concurrent load and measure origin requests.
Add single-flight with a loader timeout and test loader failure.
Compare jitter and stale-while-revalidate against a strict freshness requirement.
Interview questions and how to answer them
How do you stop a cache stampede?
Coordinate a single loader, spread expiry, and optionally serve bounded stale data while refreshing. Combine them according to freshness and failure requirements.
What if the loader crashes while holding the lock?
Use a lease with a bounded expiry and fencing or ownership validation where stale lock holders could overwrite newer work.
Why not just increase the TTL?
A longer TTL delays the next stampede and increases staleness. It does not address synchronized expiry or a single hot key.
Answers that lose the round
- Using a distributed lock with no lease or owner check.
- Serving indefinitely stale data after refresh failures.
- Refreshing every key synchronously when only a small hot set causes the load.
- Assuming random TTL jitter removes the need for origin capacity.
FAQ
Is a cache stampede the same as a cache miss?
No. A miss is one absent value; a stampede is concurrent miss work that amplifies load on the origin.
Should all waiters receive the same error?
Usually they should receive a bounded failure and retry policy. Do not let one slow or failed loader hold requests indefinitely.
Can prewarming help?
Yes for predictable hot keys, but it needs scheduling, version awareness, and limits so the warm-up itself does not overload the origin.