Distributed systems
The Redis lock expired — but the old worker did not stop
Written and reviewed by Sahil Srivastav
PTTL lock:invoice:42
(integer) -2What this error actually means
The header shows an interactive PTTL query and its missing-key reply. It is an observation, not a dedicated Redis lock exception. Redis expires a lease key; it does not stop the process that once acquired it. A worker paused by scheduling, garbage collection or a network partition can resume after another worker has acquired the same resource.
A finite TTL helps abandoned locks stop blocking progress, but it also limits how long ownership is valid. If an operation can outlive that interval, successful acquisition at the beginning is insufficient proof of exclusive authority at the moment of a later write. A timeout cannot forcibly retract an external request already in flight.
Separate two protections. An unpredictable owner token prevents an old worker from deleting or renewing a newer worker’s Redis key. A fencing token or transactional condition at the protected resource prevents an old worker from committing stale work. The first protection does not imply the second.
Causes, most common first
- 1The lease duration is shorter than an uncontrolled pause. Work time includes scheduler delays, process pauses and network latency, not just normal handler execution. A larger TTL changes the timing but cannot bound every possible pause. Resume behaviour must remain safe after ownership is lost.
- 2Renewal fails but the worker continues committing effects. A timer or heartbeat is not a guarantee of execution. The renewal may be delayed, rejected or have an ambiguous network result. Continuing as though the original acquisition remains valid creates overlapping workers.
- 3Unlock deletes by key without checking the owner. Worker A’s lease expires, B acquires the key, and A later executes DEL. That removes B’s lease. Ownership comparison and deletion must be one atomic Redis operation; a separate GET followed by DEL has its own race.
- 4The protected system accepts stale operations. A database or external API may have no record of the current lease generation. Even a perfectly implemented Redis compare-and-delete cannot prevent an old worker from issuing an already-authorised business write to that separate system.
When you see it
- Two workers process one resource after a long pause or slow downstream request
- A worker logs successful acquisition but finds PTTL negative before finishing
- Unlock code occasionally removes a lock acquired by another worker
- Increasing the TTL reduces incidents but they return during longer stalls
How to diagnose it
Step 1
Build a per-resource ownership timeline
Record acquisition attempt, successful token, renewal result, work start and final side-effect time for each worker. Use monotonic elapsed durations locally and account for cross-host clock differences. The proof is overlapping authority or a stale commit, not merely a missing key after completion.
Step 2
Inspect current lease state without assuming a snapshot
PTTL reports remaining milliseconds, -1 for no expiry and -2 for no key. GET can show the current owner token. These are separate observations and can race with expiry or acquisition; they are diagnostic evidence, not an ownership check for a subsequent write.
redis-cli PTTL lock:invoice:42
redis-cli GET lock:invoice:42Step 3
Review acquisition, renewal and release together
Check that acquisition sets NX and expiry atomically, renewal compares the owner atomically, and release does the same. Confirm that application code stops starting new protected effects after lease loss. Look beyond the lock library into callers that ignore its failure result.
Step 4
Reproduce a pause in an isolated two-worker test
Pause A after acquisition, wait beyond its TTL, allow B to acquire and commit, then resume A. Assert that A cannot remove B’s key and cannot overwrite B’s protected result. Passing only the first assertion proves token-safe release, not complete business safety.
The fix
Use a unique owner token for each acquisition attempt and set NX plus a finite expiry atomically. Release with an atomic compare-and-delete script, as shown below. The caller must supply the exact token it acquired; the illustrative key and token must not be reused across unrelated attempts.
If renewal is needed, compare the token and extend the expiry atomically. Treat an expired lease, rejected renewal or unresolved ownership as loss of permission to begin further protected work. This improves lifecycle handling but cannot cancel a database write or HTTP request that has already escaped the worker.
Enforce the real invariant at the destination. A fenced resource records the highest accepted monotonically increasing generation and rejects older generations atomically with the mutation. The generation allocator must retain its ordering guarantees across the failures you tolerate; a random owner token is not a fencing sequence, and an asynchronously replicated counter is not automatically sufficient.
Where a database already owns the business state, a transaction, version condition or unique idempotency record can be a clearer correctness boundary. For an external API, use its idempotency contract where available. Do not claim a Redis lease makes a non-idempotent external effect exactly once.
-- Atomic release only; this does not fence external writes.
-- KEYS[1] is the lock key; ARGV[1] is this attempt’s owner token.
if redis.call("GET", KEYS[1]) == ARGV[1] then
return redis.call("DEL", KEYS[1])
end
return 0How to stop it coming back
- Include long process pauses, renewal failures and ambiguous network outcomes in concurrency checks.
- Monitor lease loss separately from work failure, and require every protected operation to define its stale-worker behaviour.
- Review the failure guarantees of both the lease store and the destination; correctness spans both systems.
FAQ
Can I just set a much longer TTL?
It reduces some overlaps while delaying recovery after a crashed owner. It does not establish a bound on arbitrary pauses. Choose a practical TTL, then make stale-worker behaviour safe independently.
Does checking PTTL immediately before writing prove ownership?
No. The worker can pause after the check or the request can arrive after expiry. Enforce a condition atomically at the resource being mutated if stale writes must be rejected.
Does compare-and-delete make the lock fully safe?
It prevents an old owner from deleting a successor’s key. That is necessary lifecycle protection, but it does not prevent the old owner writing to a database or external service. The protected invariant needs its own enforcement.
Related
Other errors engineers hit next to this one
- EADDRINUSE: address already in use
- ERR_HTTP_HEADERS_SENT
- Event loop blocked by synchronous work
- pg client already connected or released twice
- Sequelize / Knex pool acquire timeout
- ERR_STREAM_PREMATURE_CLOSE during an upload
- Process exits before asynchronous writes finish
- ERR_MODULE_NOT_FOUND during ESM/CommonJS migration