Distributed systems
Retry storm and thundering herd
Written and reviewed by Sahil Srivastav
upstream request failed: 503 Service Unavailable
retrying in 1s
retrying in 1s
retrying in 1sWhat this error actually means
A retry storm happens when many callers observe one failure and retry on the same schedule. If every client waits exactly one second, they wake together, overload the recovering dependency, fail together, and repeat. The retry traffic can be larger than the original traffic.
Backoff reduces the average attempt rate; jitter spreads attempts over time. Neither makes a non-idempotent operation safe. A retry policy also needs a total deadline and attempt budget, otherwise a request that is already doomed keeps occupying its caller.
The first diagnostic is to compare request rate with attempt rate. If attempts per logical operation rise during the incident, the client is amplifying the outage.
Causes, most common first
- 1Fixed retry delay. Every instance wakes on the same second.
- 2Retries at multiple layers. SDK, service, queue consumer, and ingress each retry the same operation.
- 3No cap or budget. A request can spend its whole lifetime retrying.
- 4Retrying permanent errors. Validation and authentication failures create useless load.
When you see it
- Attempt rate increases after 5xx responses
- Dependency recovery is followed by another synchronized spike
- Many clients log identical retry timestamps
- Queues and connection pools fill while user traffic is flat
How to diagnose it
Step 1
Compare logical requests with attempts
Add an operation id and count attempts per id.
awk '$0 ~ /attempt=/ {print $0}' app.log | tail -100Step 2
Plot retry timing
A narrow histogram around one-second boundaries proves synchronization.
rg 'retrying in' app.log | awk '{print $1}' | sort | uniq -cStep 3
Trace retry layers
Inspect client, service, queue, and proxy configuration for overlapping policies.
The fix
Retry only transient failures and idempotent operations.
Use capped exponential backoff with full or equal jitter and a per-operation deadline.
Set a retry budget tied to successful request volume; shed retries when the budget is exhausted.
Choose one owner for retries per call path and pass attempt metadata downstream.
Use circuit breaking and load shedding so recovery traffic stays below dependency capacity.
delay = min(base * 2 ** attempt, cap)
await sleep(random.uniform(0, delay))How to stop it coming back
- Graph attempts per logical request
- Test recovery under synchronized clients
- Document retry ownership
- Propagate deadlines
- Alert on retry-budget consumption
FAQ
Is exponential backoff enough?
No. Without jitter, exponential schedules still synchronize; without a cap and deadline, they can keep traffic alive too long.
Should every 500 be retried?
Only when the operation and response semantics make a retry safe. A 500 after a committed write leaves outcome unknown.
Does a circuit breaker replace retries?
No. It prevents calls while a dependency is failing; bounded retries can still handle isolated transient failures.
Related
Other errors engineers hit next to this one
- Cache stampede — the same expensive value built many times concurrently
- Lock convoy — throughput collapses as threads are added, with no deadlock
- ReadWriteLock writer blocked indefinitely behind a stream of readers
- Worker loop never sees the stop flag and runs forever
- psycopg2.InterfaceError: connection already closed
- RecursionError: maximum recursion depth exceeded
- MemoryError: unable to allocate array
- QueuePool limit of size 5 overflow 10 reached, connection timed out