Performance
Cache invalidation: interview questions and practical design
Cache invalidation removes or versions cached data when the source changes so readers do not observe stale state beyond the allowed policy.
Written and reviewed by Sahil Srivastav
What it actually is
Cache invalidation removes or versions cached data when the source changes so readers do not observe stale state beyond the allowed policy.
A cache can return an old value after a successful write even when the write path is correct. Invalidating only one of several mutation paths creates intermittent stale reads.
The useful interview answer is precise about the boundary: Deleting on write avoids calculating a complete cached representation but allows a miss race; updating can preserve hit rate while risking partial or out-of-order data.
Why it matters in production
A cache can return an old value after a successful write even when the write path is correct. Invalidating only one of several mutation paths creates intermittent stale reads.
Invalidation is a distributed coordination problem: failures, ordering, and missed messages must have a recovery path.
How it works
Delete versus update
Deleting on write avoids calculating a complete cached representation but allows a miss race; updating can preserve hit rate while risking partial or out-of-order data.
Versioned keys
Include a generation or content version in the key and advance it on change. Old entries can expire naturally, which avoids a large delete but needs a source of truth for the version.
Event-driven invalidation
Publish change events after durable commit and let consumers invalidate. At-least-once delivery requires idempotent invalidation and reconciliation for missed events.
Implementing it
List every mutation path for one resource and choose one invalidation owner.
Simulate an invalidation message arriving before a cache update and after a consumer restart.
Add a manual or scheduled rebuild path for missed invalidations.
Interview questions and how to answer them
Why is invalidation hard?
The source write and cache operation are separate effects. Crashes, reordering, and concurrent writes can leave a cache stale, so the design needs versioning, retries, and repair.
Delete or update on write?
Delete is simpler when the next read can rebuild safely; update can reduce misses but must handle concurrent writers and representation construction failures.
How do you recover from missed invalidations?
Use durable change logs, version checks, periodic reconciliation, or bounded TTL as a backstop. Make the recovery observable and testable.
Answers that lose the round
- Invalidating only the endpoint that first revealed the stale data.
- Publishing an invalidation before the database commit is durable.
- Assuming cache deletion is atomic with the source write.
- Using a TTL as the only answer when users require immediate correctness.
FAQ
Can a database trigger invalidate a cache?
It can publish a change signal, but delivery, retries, and cache ownership still need design. A trigger does not make two systems atomic.
Does write-through solve invalidation?
It coordinates one write path, but other writers, deletes, and failures can still make the cache stale.
What freshness guarantee should I state?
Name the maximum stale window or read-your-own-write requirement. “Eventually consistent cache” is too vague to evaluate.