Distributed systems

Read-your-own-write failing on a replica

Written and reviewed by Sahil Srivastav

Distributed systemsReplicationConsistency
POST /profile 201 Created
GET /profile 200 OK {"displayName":"old"}  replica=read-2

What this error actually means

A write acknowledged by the primary does not imply every replica has replayed it. A subsequent read routed by a load balancer can reach a replica whose replay position is older than the write.

The user-visible guarantee is read-your-writes consistency: after a client observes success, its later reads must include that write. It is stronger than eventual consistency and weaker than routing every request to the primary.

The fix requires carrying evidence of the write forward, such as a commit LSN, session token, sticky primary route, or a bounded wait on a replica.

Causes, most common first

  1. 1Replica replay lag. WAL has not been applied on the selected replica.
  2. 2Load balancer ignores session context. The read is sent to any backend.
  3. 3Cache returns pre-write data. An application or CDN cache outlives the write.

When you see it

  • A refresh immediately after POST shows old data
  • The result changes after a few hundred milliseconds
  • Only requests routed to replicas fail
  • Tests pass with one database but fail in production

How to diagnose it

Step 1

Log backend and replay position

Compare the write LSN with replica replay LSN.

SELECT pg_current_wal_lsn();
SELECT pg_last_wal_replay_lsn(), now()-pg_last_xact_replay_timestamp();

Step 2

Trace route and cache headers

Record database host, replica identity, cache key, and age for both calls.

Step 3

Measure the consistency window

Repeat reads at intervals and report the distribution, not just the average.

The fix

Route the post-write session to primary for a bounded window.

Return a commit position and make the reader wait until its replica reaches it.

Invalidate or version caches on the write path.

Use an explicit stale-read policy for pages where the guarantee is unnecessary.

Do not hide unbounded lag with an infinite wait; fail or route primary after a deadline.

token = writer.commit_token()
return reader.get('/profile', headers={'X-Min-Commit': token})

How to stop it coming back

  • Track replica replay lag and oldest stale read
  • Test immediate read-after-write under load
  • Propagate consistency tokens through clients
  • Separate cache and database freshness metrics
  • Document endpoint consistency guarantees

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Is replication broken?

Not necessarily. Eventual replication can be healthy while failing a stronger read-your-writes contract.

Should every read use primary?

Only where freshness requires it. Deliberate routing preserves replica capacity for tolerant reads.

Can a sleep fix tests?

It hides a timing race and fails under different lag. Use a commit token or explicit consistency wait.

Related

Other errors engineers hit next to this one

Full error and symptom index →