Distributed systems

Redis OOM command not allowed — find why memory cannot be reclaimed

Written and reviewed by Sahil Srivastav

RedisMemory policyWrite rejection
OOM command not allowed when used memory > 'maxmemory'.

What this error actually means

This reply comes from Redis while the server is still running. A command that may grow memory has been rejected because the configured memory policy cannot make sufficient room. It is different from a kernel OOM kill, where the process disappears and clients normally see connection failures rather than a Redis error reply.

Start with the data contract. Evicting a derived cache entry can be acceptable because the application can rebuild it. Evicting an idempotency record, queue entry or lock can change business behaviour. Choosing an eviction policy is therefore a correctness decision before it is a memory-tuning decision.

The useful incident question is which input or retained population exceeded the intended budget. A key count alone is insufficient: one growing collection or large value can dominate memory. A TTL alone is also insufficient if the creation rate multiplied by retention time already exceeds capacity.

Causes, most common first

  1. 1A noeviction instance reaches its intended ceiling. Redis is preserving existing data and rejecting growth. This may be the correct failure mode for non-disposable state. Investigate unbounded retention or input size before replacing rejection with silent loss of existing keys.
  2. 2A volatile policy has too few eligible keys. Policies restricted to expiring keys cannot reclaim non-expiring entries. A missing TTL on the largest namespace can therefore leave the server unable to free enough memory even though eviction was configured.
  3. 3One collection or request grows much larger than expected. An ever-growing list, hash or cached report can overwhelm a budget sized from average values. Look for the owner and supported maximum size; deleting an arbitrary neighbouring cache key does not bound the next allocation.
  4. 4Capacity changed without an application-level retention review. A smaller instance, longer TTL or additional workload can make the previous configuration inadequate. Mixed cache and durable-state usage complicates recovery because the same eviction policy applies across keys in the instance.

When you see it

  • Writes fail with OOM while reads of existing data still succeed
  • Errors start after a bulk import, tenant growth or TTL configuration change
  • The server remains reachable and the process has not been restarted
  • Eviction metrics stay flat despite apparent memory pressure

How to diagnose it

Step 1

Read memory and policy together

Commands here assume redis-cli is already directed at the affected instance using your normal connection options. Compare used_memory, maxmemory and mem_not_counted_for_evict; Redis accounting is not identical to process RSS. CONFIG access may require an operator on managed services.

redis-cli INFO memory
redis-cli CONFIG GET maxmemory
redis-cli CONFIG GET maxmemory-policy

Step 2

Compare errors, evictions and expirations over time

Take two samples rather than interpreting lifetime totals as a current rate. Correlate the increase in rejected commands with a deployment or ingestion burst. A server restart resets many counters, so include uptime in the timeline.

redis-cli INFO stats
redis-cli INFO commandstats
redis-cli INFO server

Step 3

Check the suspected namespace without KEYS

Sample known keys from application logs or inventory. TTL returns -1 for a present key without expiry and -2 for an absent key. MEMORY USAGE estimates the key’s memory cost; compare representative small and large tenants.

redis-cli TYPE cache:report:42
redis-cli TTL cache:report:42
redis-cli MEMORY USAGE cache:report:42

Step 4

Relate occupancy to producer behaviour

Inspect key creation rate, retained collection lengths and TTL assignment on the actual write path. Missing expiration on an exceptional path is easy to overlook. If you need a keyspace scan, pace SCAN calls and avoid adding a blocking KEYS request during the incident.

The fix

Stop or bound the producer responsible for runaway growth. Limit collection size, maximum value size and supported cache population. Where expiry is part of the contract, set it atomically with creation so a crash cannot leave a supposedly temporary key immortal. The sample is for a disposable cache value with a deliberately chosen retention period.

If the instance stores only reconstructible cache entries, choose an eviction policy appropriate to observed access patterns and provision enough capacity for the useful working set. If it stores non-disposable state, retain rejection semantics until safe retention or additional capacity is established. Separate workloads when their loss contracts differ.

Budget process and persistence overhead beyond maxmemory before increasing it. A configuration that removes write errors by consuming all host memory can turn a controlled rejection into an abrupt process kill. Measure under replication and persistence activity as well as idle steady state.

Treat failed writes as failures in the client. In particular, do not continue a protected operation if storing its required idempotency or coordination state was rejected. Verify recovery with the same command type and input size that failed, then watch memory return to a bounded trajectory.

# Disposable cache entry; 300 seconds is an application choice.
redis-cli SET cache:report:42 '{"status":"ready"}' EX 300

How to stop it coming back

  • Assign an owner, size bound and retention policy to every key namespace.
  • Alert on rejected writes as well as occupancy; a reachable Redis server can still be unable to accept essential state.
  • Load-test large tenants and persistence activity, and keep disposable cache data separate from correctness-critical records.

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Should I set maxmemory to zero?

Removing the Redis limit transfers the boundary to the host or container. It does not make an unbounded workload bounded. Fix growth and establish a memory budget before changing the enforcement limit.

Why does eviction not happen with volatile-lru?

Only keys with an expiry are eligible. Inspect the TTLs of the namespace consuming memory; a few expiring small keys cannot compensate for a large non-expiring collection.

Is this the same as OOMKilled?

No. This is a Redis protocol error rejecting a command. OOMKilled describes process termination by the runtime or kernel. Preserve the original symptom because the evidence and recovery steps differ.

Related

Other errors engineers hit next to this one

Full error and symptom index →