Performance

Caching strategies: interview questions and practical design

Caching strategies keep reusable results closer to callers or computation so latency and backend load fall within an explicit freshness and invalidation policy.

Written and reviewed by Sahil Srivastav

PerformanceBackend engineeringInterview guide

What it actually is

Caching strategies keep reusable results closer to callers or computation so latency and backend load fall within an explicit freshness and invalidation policy.

A cache changes the source of truth seen by readers. The design is only correct when stale data, misses, evictions, and failures have defined behaviour.

The useful interview answer is precise about the boundary: The application reads the cache, loads on a miss, then stores the result. It is simple and explicit, but concurrent misses can stampede and invalidation follows every write path.

Why it matters in production

A cache changes the source of truth seen by readers. The design is only correct when stale data, misses, evictions, and failures have defined behaviour.

Caching the wrong key or scope can leak tenant data or serve a result under the wrong permissions, making correctness more important than hit rate.

How it works

Cache-aside

The application reads the cache, loads on a miss, then stores the result. It is simple and explicit, but concurrent misses can stampede and invalidation follows every write path.

Read and write policies

Write-through updates the cache as part of the write path; write-back buffers writes and trades durability and complexity for throughput. Choose based on loss and freshness tolerance.

Key and scope

Include all inputs that change the result, tenant and authorisation scope where needed, and a schema version. A correct TTL cannot repair an incomplete key.

Implementing it

Choose a cache policy for a read-heavy resource and state stale-on-error behaviour.

Add bounded expiry, negative-cache handling, and metrics for hits, misses, and evictions.

Test a permission change and a tenant boundary against cached data.

Interview questions and how to answer them

Cache-aside or write-through?

Cache-aside keeps application control and works well for read-heavy data; write-through simplifies freshness for writes but couples write latency and cache availability. State the workload and failure trade-off.

What happens when the cache is down?

Choose fail-open to the source with load protection, stale serving, or fail-closed for sensitive data. Bound the fallback so the cache outage does not become a database outage.

How do you measure a cache?

Track hit ratio by key family, miss latency, load amplification, eviction, freshness, and fallback errors—not hit ratio alone.

Answers that lose the round

  • Adding a cache before measuring the slow path.
  • Caching an unauthorised response or omitting a filter from the key.
  • Treating TTL as a complete invalidation strategy.
  • Failing closed on every cache outage when the source could safely serve stale data.

Practise in a real repository

Explaining a concept and enforcing it in code are different skills, and machine coding rounds test the second. Gronex ships broken backend repositories whose test suites assert the invariant rather than the happy path.

FAQ

Is a longer TTL better?

Only if staleness is acceptable. A long TTL can reduce load while making corrections invisible; use invalidation or versioned keys when freshness matters.

Can I cache personalised responses?

Yes only with an explicit identity and permission scope in the key or a safe private cache boundary. Shared caches must never mix users or tenants.

What is the first caching question in an interview?

Ask what may be stale, for how long, and what happens after a write. Those answers determine policy more than the cache product.

More backend concepts