Distributed systems

Distributed lock lease expiring mid-operation

Written and reviewed by Sahil Srivastav

Distributed systemsLockingCorrectness
IllegalStateException: lock lease expired; operation may no longer be exclusive

What this error actually means

A lease is time-based ownership. When its expiry passes, another worker may acquire the lock even if the first worker is paused, partitioned, or still executing. The old worker can resume and write unless the protected resource rejects stale ownership.

Renewal reduces accidental expiry but cannot make a lease proof of exclusivity: a stop-the-world pause or network delay can exceed the renewal interval. The robust mechanism is a monotonically increasing fencing token returned on acquisition and checked by the storage layer on every write.

A lock service saying “held” is not enough. Correctness belongs at the resource that can reject an old token.

Causes, most common first

  1. 1Critical section exceeds lease. The operation duration is longer than the validity window.
  2. 2Renewal thread is paused. GC, event-loop blocking, or CPU starvation prevents heartbeats.
  3. 3Clock or network delay. The holder cannot reach the lock service before expiry.
  4. 4No fencing at the resource. An expired holder’s write is accepted anyway.

When you see it

  • Two workers update the same record during a pause
  • Lock metrics show normal acquisition but duplicate side effects
  • Failures correlate with GC, CPU throttling, or network partitions
  • A watchdog renews successfully until one long operation blocks it

How to diagnose it

Step 1

Correlate token, lease, and write times

Log acquisition token, expiry, renewal, and every protected write.

rg 'lock=(acquired|renewed|expired)|fence=' app.log | sort -k1,2

Step 2

Inspect pause sources

Check GC and scheduler delay around the exact expiry.

jcmd <pid> GC.heap_info
cat /sys/fs/cgroup/cpu.stat | rg throttled

Step 3

Find stale writes

Query records where a lower token follows a higher token.

The fix

Add fencing tokens and reject writes with a token lower than the resource’s last accepted token.

Renew before half the lease, with a bounded operation deadline and explicit loss handling.

Stop work immediately when renewal fails; do not continue based on a cached “locked” flag.

Make side effects idempotent and reconcile operations interrupted by expiry.

Use a transactional database row or queue ownership when the resource already provides it.

token = lock.acquire()
try:
    store.update(key, value, fence_token=token)
finally:
    lock.release(token)

How to stop it coming back

  • Inject long pauses while holding locks
  • Monitor renewal latency and expiry margin
  • Persist fencing tokens with protected state
  • Review every lock for an owner-side timeout
  • Never claim mutual exclusion without resource enforcement

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Can I just increase the lease?

It reduces frequency but increases recovery time and cannot cover an unbounded pause. Fencing is the correctness boundary.

Does Redlock solve stale writers?

A lock algorithm does not automatically fence an external database or API. The resource still needs ownership validation.

Is renewal enough?

No. Renewal can pause or fail. Treat loss of renewal as loss of ownership.

Related

Other errors engineers hit next to this one

Full error and symptom index →