Performance

Read replica vs cache

PerformanceDatabasesDecision guide

Short answer

Use a read replica when you need more database query capacity and can tolerate its replication lag. Use a cache when repeated reads are predictable and stale or missing data is acceptable; a cache reduces work only when hit rate, invalidation, and stampede control are designed explicitly.

Written and reviewed by Sahil Srivastav

What each one actually is

A read replica replays database changes and answers ordinary queries from a full or partial copy. It preserves joins and filtering, but it is behind the writer by a measurable amount.

A cache stores selected results or objects outside the database, usually by key and TTL. It is much faster for hits, but entries can be absent, stale, malformed, or evicted at any time.

Neither fixes a bad query or an overloaded writer. A replica adds read capacity; a cache avoids work. Measure which resource is actually saturated before choosing.

Side by side

 Read replicaCache
FreshnessReplication lag; usually recent but not currentTTL or invalidation policy; can be arbitrarily stale
Query flexibilitySupports database queries and joinsFast only for designed keys or cached shapes
Miss behaviourStill serves from replica unless it failsFalls through to the database
Write visibilityRead-your-write can fail on replicaInvalidation or write-through must be handled
Capacity gainedMore query CPU and read connectionsFewer repeated queries reach the database
Failure modeLag, replay failure, promotion, stale readsEviction, stampede, stale or missing entries
Consistency contractDatabase semantics plus replica lagApplication-defined freshness and invalidation
Best access patternDiverse reads and complex filtersHot, repeated, expensive lookups

Choose Read replica when

  • Read traffic has diverse query shapes that are expensive to cache
  • The database is read-bound and consumers can tolerate measured lag
  • A cache miss would cause a costly stampede
  • You need full relational semantics for the read workload

Choose Cache when

  • The same objects or responses are requested repeatedly
  • A bounded stale window is acceptable
  • The result can be rebuilt from an authoritative source
  • You can define invalidation, TTL jitter, and bounded refresh concurrency

The trade-off in detail

A replica often looks like a cache because it serves old data, but its contract is different: it has a complete queryable state and lag is tied to replication. Route a read-after-write to the primary or use a session consistency rule when users must see their own change.

A cache can make a database appear healthy until expiry aligns across keys. Add jitter to TTLs, coalesce concurrent misses, and cap refresh work; otherwise the cache turns a quiet period into a thundering herd.

Adding a replica increases connection and replication load. Adding a cache increases invalidation and memory cost. Profile query CPU, network, hit rate, lag, and correctness errors before scaling either layer.

Things that are commonly said and are wrong

  • “A replica is a cache.” A replica is a database copy with replication semantics; it cannot be safely discarded without considering failover and lag.
  • “Cache invalidation is just deleting a key.” Multiple keys, derived lists, and races can leave stale data; versioned keys or write-through ownership can simplify the contract.
  • “Replicas solve write bottlenecks.” Writes still go to the primary in the usual topology, and indexes or lock contention may remain the limiting resource.

Decide it in a real repository

Choosing correctly on a whiteboard and enforcing the choice in code are different skills. Gronex ships broken backend repositories whose tests assert the invariant, not the happy path.

FAQ

Which should I add first?

If the database is read-CPU bound and queries are diverse, a replica is usually the more faithful scale-out. If a small number of hot results dominate and bounded staleness is acceptable, cache those results first.

How do I guarantee read-your-write with replicas?

Keep the user on the primary for a short window, wait until a replica reaches a write position, or route consistency-sensitive requests by session. A random replica cannot guarantee it.

Should cached writes update the cache?

They may, but only with a clear source of truth and failure policy. Write-through or versioned invalidation can work; silently updating a cache while the database write fails creates a false result.

Other decisions engineers weigh