Distributed systems
Redis stalls during KEYS — one administrative scan delays unrelated requests
Written and reviewed by Sahil Srivastav
KEYS *What this error actually means
The header shows the problematic command, not an error reply. KEYS enumerates matching keys in one operation. On a large keyspace, that work can occupy the command-execution path long enough to delay unrelated requests. Applications may report timeouts even though the requested GET or SET is individually cheap.
The symptom often looks like a network incident because many callers stall together and then recover. The decisive evidence is a slow enumeration or other expensive command at the start of that interval. Adding client connections does not make the server finish the blocking command sooner.
Redis’s blocked_clients metric has a narrower meaning than everyday speech: it counts clients waiting on blocking operations such as blocking list reads. Clients queued behind an expensive KEYS execution need not make that metric spike. A low blocked_clients value therefore does not rule out this latency mechanism.
Causes, most common first
- 1An application request enumerates a namespace with KEYS. The code uses key names as an implicit index and scans to discover all members. A small development database hides the cost. Production latency then scales with the keyspace rather than the number of results the user actually needs.
- 2A cleanup or monitoring job performs full scans. An administrative script can share the same Redis instance as latency-sensitive traffic. The fact that the job is outside the application request path does not isolate its server execution cost.
- 3A huge response adds transfer and client-processing cost. Enumeration is only part of end-to-end latency. Serialising, transmitting and consuming a large key list can also fill buffers and overload the caller. A short server execution time does not guarantee a cheap response.
- 4Another expensive command is being mistaken for KEYS. Large scripts, collection operations or oversized values can produce a similar stall. Preserve command-level evidence before changing enumeration code; the common symptom alone does not identify the command responsible.
When you see it
- Unrelated cache lookups time out during an administration or cleanup job
- Latency spikes grow with keyspace size despite unchanged application request volume
- A diagnostic dashboard refresh or wildcard deletion script triggers the same stall
- SLOWLOG shows KEYS near the start of a broad latency incident
How to diagnose it
Step 1
Inspect recent slow commands after responsiveness returns
SLOWLOG records server execution time, excluding network I/O. Match timestamps and command names to the incident. Entries can contain key names or arguments, so handle output according to the data it contains. An empty log may reflect its threshold or retention.
redis-cli SLOWLOG GET 20
redis-cli CONFIG GET slowlog-log-slower-than
redis-cli CONFIG GET slowlog-max-lenStep 2
Compare command counters across the incident
INFO commandstats can show KEYS call counts and execution totals. Use deltas over a known window and correlate with the job schedule; cumulative totals alone cannot identify the most recent spike.
redis-cli INFO commandstats
redis-cli INFO clientsStep 3
Measure round-trip latency separately
redis-cli --latency repeatedly probes response time until interrupted. Run it from the application network during a controlled reproduction, then compare with server execution measurements. Do not launch a large KEYS scan on production just to demonstrate the problem.
redis-cli --latencyStep 4
Locate the producer of enumeration work
Search application and maintenance code for KEYS and namespace-wide discovery. Check client names and scheduled-job times. Avoid using a global MONITOR stream as the default response to an already overloaded server; targeted evidence is usually enough.
The fix
Remove unbounded key discovery from latency-sensitive requests. If the product needs a list of an account’s objects, maintain an explicit bounded index or query the authoritative store. An index also needs lifecycle management so deleted or expired objects do not leave stale membership forever.
For maintenance that genuinely needs enumeration, use SCAN in paced increments. Start with cursor 0 and continue using each returned cursor until it returns 0 again. An empty result batch with a nonzero cursor is not completion. COUNT is a work hint, not a guaranteed page size or a hard time limit.
Design the scan consumer for duplicates and concurrent changes. SCAN is not a stable snapshot of a changing database. A cleanup action should be idempotent, and a reporting workflow that needs an exact point-in-time inventory needs a different data source or coordinated snapshot procedure.
Keep the entire pipeline bounded: scanned keys, follow-up reads, result buffering and writes. Replacing one KEYS call with an unthrottled loop of SCAN plus thousands of parallel GETs can still overload the instance. In Cluster, plan enumeration across the intended nodes; one node-local scan is not automatically a complete cluster inventory.
# One incremental inspection step; continue with the returned cursor.
redis-cli SCAN 0 MATCH 'cache:*' COUNT 100
# COUNT is a hint. Pace subsequent calls and tolerate duplicate keys.How to stop it coming back
- Keep namespace enumeration out of ordinary request handlers and review administrative jobs against realistic key counts.
- Name application and maintenance connections so command-level incidents can be attributed to their owner.
- Alert on server execution latency and client round-trip latency separately, and retain enough slow-log history to cover incident detection delay.
FAQ
Is KEYS with a narrow pattern cheap?
A selective pattern can still require substantial enumeration of the keyspace. Do not infer cost from the small number of returned names. Use measured behaviour and an incremental or indexed access path for growing production data.
Does SCAN return each key exactly once?
No. The consumer must tolerate duplicates, and changes during iteration prevent general snapshot semantics. Make maintenance actions safe to repeat and avoid using one scan as proof of a consistent inventory.
Will increasing client timeouts solve the problem?
It can hide the immediate exception while requests and memory accumulate upstream. Reduce or isolate the expensive work and bound callers; timeout tuning should follow the measured latency budget.
Related
Other errors engineers hit next to this one
- QueuePool limit of size 5 overflow 10 reached, connection timed out
- DetachedInstanceError: instance is not bound to a Session
- RuntimeError: Event loop is closed
- Task was destroyed but it is pending!
- Executing <Handle ...> took 2.418 seconds (blocked event loop)
- SettingWithCopyWarning: A value is trying to be set on a copy of a slice
- celery.exceptions.WorkerLostError: Worker exited prematurely
- requests.exceptions.ReadTimeout: HTTPSConnectionPool read timed out