Distributed systems

Clock skew breaking timestamp ordering or token expiry

Written and reviewed by Sahil Srivastav

Distributed systemsTimeCorrectness
token expired at 10:00:00.100 on node-a but node-b reports current time 09:59:59.800

What this error actually means

Wall clocks are estimates maintained by synchronisation protocols. They can jump forward or backward, differ between hosts, and have uncertainty during suspend, VM migration, or an NTP outage. A timestamp from node A is therefore not a proof that an event happened after one from node B.

Expiry calculations are especially fragile when one node issues a token and another validates it. A fast validator rejects valid work; a slow validator accepts expired work. Subtracting two readings from different machines has the same problem.

Use monotonic clocks for durations, a trusted authority for expiry, and logical or database ordering for event sequences.

Causes, most common first

  1. 1NTP offset or failed synchronisation. Hosts disagree about wall time.
  2. 2Wall clock used for elapsed duration. A clock correction makes a timeout negative or huge.
  3. 3Client-supplied timestamps trusted. Untrusted clocks influence expiry or ordering.
  4. 4Milliseconds versus seconds mismatch. Unit errors mimic skew.

When you see it

  • Only some hosts reject tokens
  • Events appear out of order despite increasing timestamps
  • Failures begin after VM resume or NTP alarms
  • Retries succeed when routed to another instance

How to diagnose it

Step 1

Measure offset on every host

Compare system time to the configured time source and record stratum.

chronyc tracking
timedatectl status

Step 2

Log monotonic and wall readings

A monotonic delta proves whether the local clock jumped during an operation.

Step 3

Inspect token issuer and validator

Compare their clock offsets and units, then test a token near its boundary.

The fix

Use a monotonic clock for local durations and deadlines.

Have the issuer include server time and expiry, or validate against one trusted service.

Use database sequence, log offset, or hybrid logical clocks for ordering.

Allow a narrowly justified skew tolerance only for authentication, with bounded maximum lifetime.

Reject impossible units and timestamps at the protocol boundary.

started = time.monotonic()
while time.monotonic() - started < timeout_s:
    work()

How to stop it coming back

  • Alert on NTP offset and unsynchronised hosts
  • Never compare wall timestamps as causal proof
  • Test clock jumps and suspend/resume
  • Document timestamp units
  • Keep expiry validation in one library

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Is NTP enough for ordering?

No. NTP reduces offset but does not provide a causal ordering guarantee.

Why not use UTC everywhere?

UTC is the right representation for wall time, but it remains an adjustable wall clock and is unsuitable for elapsed durations.

How much skew tolerance is safe?

Only the measured synchronisation bound plus operational margin; a large tolerance weakens expiry.

Related

Other errors engineers hit next to this one

Full error and symptom index →