Distributed systems

Replication lag serving stale data after a write

Written and reviewed by Sahil Srivastav

Distributed systemsReplicationOperations
replica replay_lag=47s; read returned WAL position 0/16A2 before required 0/1A90

What this error actually means

Replication is a pipeline: the primary generates log, a sender transmits it, and the replica receiver, writer, and replay process apply it. “Replica lag” is not one number. A replica can have received WAL but not replayed it, or the sender can be blocked before transmission.

Serving from a lagging replica is correct only when the endpoint permits stale data. After a write, stale reads violate user expectations and can trigger incorrect decisions such as reordering, duplicate creation, or inventory display errors.

The safe response is to measure the stage, protect freshness-sensitive reads, and apply backpressure when lag threatens the service’s contract.

Causes, most common first

  1. 1Replica I/O or replay saturation. The receiver cannot write or apply WAL as quickly as it arrives.
  2. 2Long-running transaction. Replay waits behind a snapshot or conflicting query.
  3. 3Network bandwidth or sender slot pressure. WAL cannot reach the replica in time.
  4. 4A burst exceeds steady-state capacity. Lag accumulates faster than it drains.

When you see it

  • Replay lag grows during write bursts
  • Replica CPU or disk is saturated
  • Read results differ by backend
  • Failover candidates are far behind the primary
  • Read-your-own-write failures correlate with lag

How to diagnose it

Step 1

Separate receive from replay lag

Compare sender, flush, and replay positions.

SELECT client_addr, sent_lsn, write_lsn, flush_lsn, replay_lsn, state FROM pg_stat_replication;
SELECT pg_last_wal_receive_lsn(), pg_last_wal_replay_lsn();

Step 2

Find replay blockers

Inspect long queries and conflicts on the replica.

SELECT pid, usename, state, now()-query_start AS age, query FROM pg_stat_activity WHERE state <> 'idle' ORDER BY age DESC;

Step 3

Check disk and network

Correlate lag with I/O latency, WAL volume, and link saturation.

The fix

Route freshness-sensitive reads to primary or use a minimum replay position token.

Remove long-running replica queries or isolate analytics on another replica.

Provision I/O and network for peak WAL rate with recovery headroom.

Throttle write bursts or queue rebuild jobs when lag crosses the SLO.

Do not promote a replica without checking its replay position and data-loss policy.

required = write_response.commit_lsn
if replica.replay_lsn < required:
    return primary.read(request)
return replica.read(request)

How to stop it coming back

  • Alert on lag seconds and WAL byte distance
  • Track stale-read rate by endpoint
  • Load-test failover and replay catch-up
  • Keep analytics off serving replicas
  • Publish consistency guarantees to API owners

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Is a read replica always eventually consistent?

Usually by design, but a broken or blocked replay pipeline can make the convergence window unbounded.

Does adding replicas reduce lag?

It can distribute reads, but it does not increase a saturated primary-to-replica link or replay disk throughput.

Can caches make this worse?

Yes. A cache can extend the stale window beyond database replay unless invalidated or versioned.

Related

Other errors engineers hit next to this one

Full error and symptom index →