Distributed systems

Retry storm and thundering herd

Written and reviewed by Sahil Srivastav

Distributed systemsResilienceRetries
upstream request failed: 503 Service Unavailable
retrying in 1s
retrying in 1s
retrying in 1s

What this error actually means

A retry storm happens when many callers observe one failure and retry on the same schedule. If every client waits exactly one second, they wake together, overload the recovering dependency, fail together, and repeat. The retry traffic can be larger than the original traffic.

Backoff reduces the average attempt rate; jitter spreads attempts over time. Neither makes a non-idempotent operation safe. A retry policy also needs a total deadline and attempt budget, otherwise a request that is already doomed keeps occupying its caller.

The first diagnostic is to compare request rate with attempt rate. If attempts per logical operation rise during the incident, the client is amplifying the outage.

Causes, most common first

  1. 1Fixed retry delay. Every instance wakes on the same second.
  2. 2Retries at multiple layers. SDK, service, queue consumer, and ingress each retry the same operation.
  3. 3No cap or budget. A request can spend its whole lifetime retrying.
  4. 4Retrying permanent errors. Validation and authentication failures create useless load.

When you see it

  • Attempt rate increases after 5xx responses
  • Dependency recovery is followed by another synchronized spike
  • Many clients log identical retry timestamps
  • Queues and connection pools fill while user traffic is flat

How to diagnose it

Step 1

Compare logical requests with attempts

Add an operation id and count attempts per id.

awk '$0 ~ /attempt=/ {print $0}' app.log | tail -100

Step 2

Plot retry timing

A narrow histogram around one-second boundaries proves synchronization.

rg 'retrying in' app.log | awk '{print $1}' | sort | uniq -c

Step 3

Trace retry layers

Inspect client, service, queue, and proxy configuration for overlapping policies.

The fix

Retry only transient failures and idempotent operations.

Use capped exponential backoff with full or equal jitter and a per-operation deadline.

Set a retry budget tied to successful request volume; shed retries when the budget is exhausted.

Choose one owner for retries per call path and pass attempt metadata downstream.

Use circuit breaking and load shedding so recovery traffic stays below dependency capacity.

delay = min(base * 2 ** attempt, cap)
await sleep(random.uniform(0, delay))

How to stop it coming back

  • Graph attempts per logical request
  • Test recovery under synchronized clients
  • Document retry ownership
  • Propagate deadlines
  • Alert on retry-budget consumption

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Is exponential backoff enough?

No. Without jitter, exponential schedules still synchronize; without a cap and deadline, they can keep traffic alive too long.

Should every 500 be retried?

Only when the operation and response semantics make a retry safe. A 500 after a committed write leaves outcome unknown.

Does a circuit breaker replace retries?

No. It prevents calls while a dependency is failing; bounded retries can still handle isolated transient failures.

Related

Other errors engineers hit next to this one

Full error and symptom index →