Distributed systems
Redis OOM command not allowed — find why memory cannot be reclaimed
Written and reviewed by Sahil Srivastav
OOM command not allowed when used memory > 'maxmemory'.What this error actually means
This reply comes from Redis while the server is still running. A command that may grow memory has been rejected because the configured memory policy cannot make sufficient room. It is different from a kernel OOM kill, where the process disappears and clients normally see connection failures rather than a Redis error reply.
Start with the data contract. Evicting a derived cache entry can be acceptable because the application can rebuild it. Evicting an idempotency record, queue entry or lock can change business behaviour. Choosing an eviction policy is therefore a correctness decision before it is a memory-tuning decision.
The useful incident question is which input or retained population exceeded the intended budget. A key count alone is insufficient: one growing collection or large value can dominate memory. A TTL alone is also insufficient if the creation rate multiplied by retention time already exceeds capacity.
Causes, most common first
- 1A noeviction instance reaches its intended ceiling. Redis is preserving existing data and rejecting growth. This may be the correct failure mode for non-disposable state. Investigate unbounded retention or input size before replacing rejection with silent loss of existing keys.
- 2A volatile policy has too few eligible keys. Policies restricted to expiring keys cannot reclaim non-expiring entries. A missing TTL on the largest namespace can therefore leave the server unable to free enough memory even though eviction was configured.
- 3One collection or request grows much larger than expected. An ever-growing list, hash or cached report can overwhelm a budget sized from average values. Look for the owner and supported maximum size; deleting an arbitrary neighbouring cache key does not bound the next allocation.
- 4Capacity changed without an application-level retention review. A smaller instance, longer TTL or additional workload can make the previous configuration inadequate. Mixed cache and durable-state usage complicates recovery because the same eviction policy applies across keys in the instance.
When you see it
- Writes fail with OOM while reads of existing data still succeed
- Errors start after a bulk import, tenant growth or TTL configuration change
- The server remains reachable and the process has not been restarted
- Eviction metrics stay flat despite apparent memory pressure
How to diagnose it
Step 1
Read memory and policy together
Commands here assume redis-cli is already directed at the affected instance using your normal connection options. Compare used_memory, maxmemory and mem_not_counted_for_evict; Redis accounting is not identical to process RSS. CONFIG access may require an operator on managed services.
redis-cli INFO memory
redis-cli CONFIG GET maxmemory
redis-cli CONFIG GET maxmemory-policyStep 2
Compare errors, evictions and expirations over time
Take two samples rather than interpreting lifetime totals as a current rate. Correlate the increase in rejected commands with a deployment or ingestion burst. A server restart resets many counters, so include uptime in the timeline.
redis-cli INFO stats
redis-cli INFO commandstats
redis-cli INFO serverStep 3
Check the suspected namespace without KEYS
Sample known keys from application logs or inventory. TTL returns -1 for a present key without expiry and -2 for an absent key. MEMORY USAGE estimates the key’s memory cost; compare representative small and large tenants.
redis-cli TYPE cache:report:42
redis-cli TTL cache:report:42
redis-cli MEMORY USAGE cache:report:42Step 4
Relate occupancy to producer behaviour
Inspect key creation rate, retained collection lengths and TTL assignment on the actual write path. Missing expiration on an exceptional path is easy to overlook. If you need a keyspace scan, pace SCAN calls and avoid adding a blocking KEYS request during the incident.
The fix
Stop or bound the producer responsible for runaway growth. Limit collection size, maximum value size and supported cache population. Where expiry is part of the contract, set it atomically with creation so a crash cannot leave a supposedly temporary key immortal. The sample is for a disposable cache value with a deliberately chosen retention period.
If the instance stores only reconstructible cache entries, choose an eviction policy appropriate to observed access patterns and provision enough capacity for the useful working set. If it stores non-disposable state, retain rejection semantics until safe retention or additional capacity is established. Separate workloads when their loss contracts differ.
Budget process and persistence overhead beyond maxmemory before increasing it. A configuration that removes write errors by consuming all host memory can turn a controlled rejection into an abrupt process kill. Measure under replication and persistence activity as well as idle steady state.
Treat failed writes as failures in the client. In particular, do not continue a protected operation if storing its required idempotency or coordination state was rejected. Verify recovery with the same command type and input size that failed, then watch memory return to a bounded trajectory.
# Disposable cache entry; 300 seconds is an application choice.
redis-cli SET cache:report:42 '{"status":"ready"}' EX 300How to stop it coming back
- Assign an owner, size bound and retention policy to every key namespace.
- Alert on rejected writes as well as occupancy; a reachable Redis server can still be unable to accept essential state.
- Load-test large tenants and persistence activity, and keep disposable cache data separate from correctness-critical records.
FAQ
Should I set maxmemory to zero?
Removing the Redis limit transfers the boundary to the host or container. It does not make an unbounded workload bounded. Fix growth and establish a memory budget before changing the enforcement limit.
Why does eviction not happen with volatile-lru?
Only keys with an expiry are eligible. Inspect the TTLs of the namespace consuming memory; a few expiring small keys cannot compensate for a large non-expiring collection.
Is this the same as OOMKilled?
No. This is a Redis protocol error rejecting a command. OOMKilled describes process termination by the runtime or kernel. Preserve the original symptom because the evidence and recovery steps differ.
Related
Other errors engineers hit next to this one
- Intermittent 502 after an idle keep-alive connection
- SSL certificate problem: unable to get local issuer certificate
- ERR_INCOMPLETE_CHUNKED_ENCODING
- Request timeouts cascading into pool exhaustion
- Connection reset by peer on a long-polling endpoint
- MISCONF Redis is configured to save RDB snapshots
- READONLY You can’t write against a read only replica
- LOADING Redis is loading the dataset in memory