Distributed systems
Distributed lock lease expiring mid-operation
Written and reviewed by Sahil Srivastav
IllegalStateException: lock lease expired; operation may no longer be exclusiveWhat this error actually means
A lease is time-based ownership. When its expiry passes, another worker may acquire the lock even if the first worker is paused, partitioned, or still executing. The old worker can resume and write unless the protected resource rejects stale ownership.
Renewal reduces accidental expiry but cannot make a lease proof of exclusivity: a stop-the-world pause or network delay can exceed the renewal interval. The robust mechanism is a monotonically increasing fencing token returned on acquisition and checked by the storage layer on every write.
A lock service saying “held” is not enough. Correctness belongs at the resource that can reject an old token.
Causes, most common first
- 1Critical section exceeds lease. The operation duration is longer than the validity window.
- 2Renewal thread is paused. GC, event-loop blocking, or CPU starvation prevents heartbeats.
- 3Clock or network delay. The holder cannot reach the lock service before expiry.
- 4No fencing at the resource. An expired holder’s write is accepted anyway.
When you see it
- Two workers update the same record during a pause
- Lock metrics show normal acquisition but duplicate side effects
- Failures correlate with GC, CPU throttling, or network partitions
- A watchdog renews successfully until one long operation blocks it
How to diagnose it
Step 1
Correlate token, lease, and write times
Log acquisition token, expiry, renewal, and every protected write.
rg 'lock=(acquired|renewed|expired)|fence=' app.log | sort -k1,2Step 2
Inspect pause sources
Check GC and scheduler delay around the exact expiry.
jcmd <pid> GC.heap_info
cat /sys/fs/cgroup/cpu.stat | rg throttledStep 3
Find stale writes
Query records where a lower token follows a higher token.
The fix
Add fencing tokens and reject writes with a token lower than the resource’s last accepted token.
Renew before half the lease, with a bounded operation deadline and explicit loss handling.
Stop work immediately when renewal fails; do not continue based on a cached “locked” flag.
Make side effects idempotent and reconcile operations interrupted by expiry.
Use a transactional database row or queue ownership when the resource already provides it.
token = lock.acquire()
try:
store.update(key, value, fence_token=token)
finally:
lock.release(token)How to stop it coming back
- Inject long pauses while holding locks
- Monitor renewal latency and expiry margin
- Persist fencing tokens with protected state
- Review every lock for an owner-side timeout
- Never claim mutual exclusion without resource enforcement
FAQ
Can I just increase the lease?
It reduces frequency but increases recovery time and cannot cover an unbounded pause. Fencing is the correctness boundary.
Does Redlock solve stale writers?
A lock algorithm does not automatically fence an external database or API. The resource still needs ownership validation.
Is renewal enough?
No. Renewal can pause or fail. Treat loss of renewal as loss of ownership.
Related
Other errors engineers hit next to this one
- Too many open files
- Out of memory: Killed process
- No space left on device despite free disk space
- Text file busy during executable replacement
- set -e script continues after a failed pipeline
- An unquoted variable turns one argument into several
- OOMKilled — container exit code 137
- CrashLoopBackOff