Distributed systems

Redis LOADING — distinguish slow recovery from a restart loop

Written and reviewed by Sahil Srivastav

Redis startupPersistence recoveryReadiness
LOADING Redis is loading the dataset in memory

What this error actually means

The server is reconstructing its in-memory dataset and cannot yet serve the requested command. Loading can happen during startup from persisted state or while a replica installs a full synchronisation. A reachable TCP port therefore does not prove Redis is ready for normal application traffic.

The first distinction is progress versus repetition. A large but steadily advancing load may be healthy recovery with an undersized readiness window. A process whose uptime repeatedly resets may never get the chance to finish. Restarting a progressing instance can turn a slow recovery into an indefinite outage.

A loading error is not evidence that the dataset should be deleted. Persistence exists to recover that state. Removing an RDB or AOF to make startup faster can discard the exact records the service needs. Establish what Redis is loading, whether it is advancing, and why it might be restarting before changing recovery behaviour.

Causes, most common first

  1. 1The persisted dataset takes longer to load than the startup budget. As data grows, a recovery path tested with a small fixture can exceed the configured probe or client timeout. The important quantity is measured restore time on the actual storage and CPU allocation, not normal request latency.
  2. 2A supervisor restarts Redis before recovery completes. A liveness check may mistake not-ready for dead, or a deployment controller may impose an unrealistic startup deadline. Each restart resets useful work, so repeated restarts can look like one extremely slow loading operation.
  3. 3Storage throughput or CPU allocation constrains replay. A large snapshot, lengthy AOF replay or contended volume can slow reconstruction. Correlate progress with resource metrics and Redis logs; increasing client retries does not make the server load bytes faster.
  4. 4The load fails and is retried from scratch. An OOM kill, persistence-file error or recurring full synchronisation can prevent completion. Look for process restarts, changing replication state and explicit load failures rather than assuming every LOADING reply is transient healthy startup.

When you see it

  • Clients connect successfully but ordinary commands return LOADING
  • Loading progress advances while application startup repeatedly times out
  • Redis uptime resets and loading starts over at regular intervals
  • A replica becomes unavailable during a full resynchronisation

How to diagnose it

Step 1

Observe loading progress on the affected instance

INFO is available during loading in normal Redis operation. Fields such as loading, loading_loaded_bytes, loading_total_bytes and loading_eta_seconds vary with the load path and version. Take repeated samples and treat ETA as an estimate, not a completion guarantee.

redis-cli INFO persistence

Step 2

Check whether uptime is continuous

Compare uptime_in_seconds between samples. A reset is evidence of a new process, not stalled progress in the same process. Correlate with the supervisor or container restart count and kernel memory events.

redis-cli INFO server

Step 3

Identify startup recovery versus replica synchronisation

INFO replication and the server logs help identify a full sync. A replica repeatedly dropping and rebuilding its upstream connection needs replication-path diagnosis; changing the application’s startup timeout cannot repair that loop.

redis-cli INFO replication

Step 4

Read the server’s recovery timeline

For a systemd deployment named redis-server.service, look for load start, successful completion, fatal file errors and shutdown messages. Use your deployment’s equivalent logs when the unit name differs. Align storage and CPU metrics to this window.

journalctl -u redis-server.service --since '30 minutes ago'

The fix

If loading progresses normally, allow recovery to complete and keep the instance out of application traffic until it can serve the required commands. Set a startup allowance from measured restore time plus headroom. Separate startup/readiness behaviour from liveness so a busy recovering process is not killed for being temporarily unavailable.

Clients should use bounded retries with backoff and jitter, respecting their request deadline. Do not queue unlimited requests in memory while Redis loads. Decide which application functions can degrade without Redis and which must fail explicitly until their dependency is ready.

If progress resets, repair the restart trigger: memory budget, supervisor settings, persistence-file issue or replication instability. Preserve persistence artifacts and logs before attempting recovery. Use supported validation and restore procedures on copies where appropriate, rather than deleting files to suppress the loading phase.

For consistently slow recovery, measure the cost of restoring the real dataset on the provisioned storage. Review retention, persistence strategy and recovery capacity. A healthy steady-state latency benchmark does not establish that recovery time meets the service’s availability objective.

How to stop it coming back

  • Rehearse restart and restore with production-scale data, and record completion time as a deployment input.
  • Alert on repeated uptime resets and loading without progress, while allowing a bounded normal recovery window.
  • Keep application queues and retry budgets finite so a recovering dependency does not exhaust its callers.
  • Validate backups by restoring them into an isolated instance, including the largest supported dataset and expected resource limits.

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Should I restart Redis when I see LOADING?

Only after diagnosing why it is not making progress. Restarting a healthy load discards its progress and can repeat the entire recovery. Check loaded bytes, uptime and logs first.

Can I use PING as the only readiness check?

Connectivity and command readiness are different. A loading server can reject ordinary commands even though the port accepts connections. Use a readiness procedure that recognises loading and verifies the operations the application actually needs.

Will a larger request timeout fix recovery?

It may stop a caller giving up too early, but it does not address a restart loop, storage bottleneck or invalid persistence file. Keep timeout changes tied to measured progress and a finite recovery budget.

Related

Other errors engineers hit next to this one

Full error and symptom index →