Distributed systems

Redis READONLY — your write reached a replica

Written and reviewed by Sahil Srivastav

Redis replicationFailoverClient routing
READONLY You can't write against a read only replica.

What this error actually means

The node that received the command currently considers itself a read-only replica. This is a role and routing failure, not an ACL permission error. The client may still authenticate, PING and read successfully, so a basic connectivity check cannot prove that a connection is suitable for writes.

Failover makes previously correct connections stale. A pool can retain a TCP connection to a former primary after that node has become a replica. Updating DNS or a configuration file does not retroactively move an existing socket. The application must use the discovery and reconnection behaviour appropriate to its Redis deployment.

Separate standalone replication, Sentinel-managed failover, Redis Cluster and managed-service endpoint semantics. They do not share one universal client configuration. The repair is to restore the intended writer route for that topology, while retaining replicas’ protective write behaviour.

Causes, most common first

  1. 1The application uses a reader endpoint for writes. A configuration variable may point at a replica address or a load balancer distributing connections across roles. Successful reads hide the error until the first state-changing command. Compare writer and reader configuration explicitly.
  2. 2A connection pool still points at the demoted primary. Long-lived sockets outlive DNS changes and role transitions. If the client does not refresh the topology or retire stale connections, individual instances can remain stuck on the wrong node after the infrastructure has completed failover.
  3. 3Discovery is configured for the wrong deployment mode. A plain single-node client does not automatically discover a Sentinel primary, and a seed address does not make a client cluster-aware. Verify the library’s chosen mode, master name or endpoint contract against the actual deployment.
  4. 4Network access permits only the obsolete node. A client may learn the new primary but fail to reach it because of routing or firewall changes. A fallback to the old connection then continues producing READONLY. Check access from the application network, not only an operator laptop.

When you see it

  • Writes begin failing immediately after a failover or maintenance event
  • Only some application instances fail because their pools retain different connections
  • PING and GET work through the same connection that rejects SET
  • Restarting the client helps temporarily, but the next topology change repeats the outage

How to diagnose it

Step 1

Ask the node for its current role

Run these against the exact endpoint receiving the rejected writes. ROLE identifies the server role; INFO replication supplies additional upstream and replication state. The protocol may still use the historical words master and slave in output.

redis-cli ROLE
redis-cli INFO replication

Step 2

Identify the client’s actual remote address

Use client-library connection logs or process socket inspection to match the failing pool to a node. Compare affected and unaffected application instances. Resolving the configured hostname now does not prove which address an existing socket uses.

Step 3

Check Sentinel discovery when Sentinel is actually used

For a local test Sentinel on port 26379 with master name mymaster, this returns its current primary address. Substitute the real Sentinel endpoint and configured name. Run from the application network and verify the returned address is reachable.

redis-cli -p 26379 SENTINEL get-master-addr-by-name mymaster

Step 4

Align the timeline with role changes

Compare Redis or provider failover events with connection creation times and the first READONLY reply. If failures precede any transition, prioritise static endpoint misconfiguration. If they persist only on older connections, investigate pool refresh and role discovery.

The fix

Send writes to the authoritative primary endpoint or use the topology-aware mode required by the deployment. Keep reader and writer client configuration distinct where read replicas are intentional. Avoid a generic TCP balancing pool that mixes primary and replica nodes for state-changing traffic.

On a role error, have the client refresh discovery and retire the stale connection according to the library’s supported failover behaviour. Reconnect with bounded backoff and jitter so all application instances do not hammer discovery simultaneously. Validate this through an actual controlled failover, not just a process restart.

Do not make replicas writable as a routing workaround. Local replica writes do not establish the application’s intended primary history and can be lost or conflict with later replication. The error is protecting you from writing to the wrong authority.

Distinguish an explicit READONLY rejection from a transport failure on an earlier command. The rejected command did not execute on that replica, but a batch or previous timed-out operation may have partially succeeded elsewhere. Replay only operations whose retry semantics are understood, using idempotency for external effects.

How to stop it coming back

  • Run failover exercises with long-lived pools and verify recovery without manually restarting every application instance.
  • Record node role, remote address and client topology mode in diagnostics without exposing authentication material.
  • Keep readiness checks aligned with the dependency role the application needs; a successful PING alone does not validate writer routing.

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Is this an ACL problem?

An ACL rejection normally uses a permission-related reply such as NOPERM. READONLY indicates the node’s replication role. Preserve the server reply instead of translating every write failure into a generic authentication error.

Why did refreshing DNS not fix it?

The existing TCP socket remains connected to its original peer. DNS affects later connection establishment. The pool must retire the stale connection and discover or resolve the correct writer when reconnecting.

Can replicas still be used for reads?

Yes, if the application accepts their consistency behaviour. Read routing and write routing have different contracts; a successful replica read does not make that node a valid destination for updates.

Related

Other errors engineers hit next to this one

Full error and symptom index →