Distributed systems
429 Too Many Requests — retry when capacity is available
Written and reviewed by Sahil Srivastav
429 Too Many RequestsWhat this error actually means
The server is applying a request-rate policy to some identity or resource associated with this request. The policy might be keyed by API token, account, route, tenant or source IP. A 429 from a shared NAT address does not prove that your individual process exceeded its own quota; conversely, many workers using one API token may collectively exceed an account limit.
Retry-After, when present, is either an integer number of seconds to wait or an HTTP date. It is not a millisecond timestamp. A valid value supplies a minimum wait before another attempt; the client still needs a total deadline, an attempt cap and a policy for operations whose side effects are uncertain.
A retry loop without shared admission control can turn one throttled call into dozens. If each worker wakes at exactly the advertised boundary, they create another burst together. A durable fix therefore coordinates the quota scope and adds delay beyond the lower bound, rather than shortening the wait in pursuit of an earlier success.
Causes, most common first
- 1Concurrency is coordinated locally but the quota is shared. Every process believes it is below the allowed rate. Their aggregate traffic exceeds the server’s bucket. Autoscaling can make the problem worse because adding instances adds independent senders without increasing the external quota.
- 2Retries or overlapping schedules amplify requests. An SDK retries internally, application code retries around it, and a job scheduler launches another run before the previous one ends. Count wire attempts per logical operation to see the multiplication.
- 3Retry-After is parsed incorrectly or ignored. Treating seconds as milliseconds creates an almost immediate retry. Treating an HTTP date as an integer produces nonsense or a zero delay. Missing and malformed values require a fallback policy, not a tight retry loop.
- 4A permanent quota or policy limit looks transient. A daily allowance, account restriction or endpoint-specific cap may not recover within this job’s lifetime. Retrying every few seconds consumes resources without changing the condition that caused rejection.
When you see it
- Several workers begin receiving 429 at the same instant despite moderate per-worker throughput.
- Failures repeat at regular intervals after a retry timer releases a backlog.
- A polling integration consumes its quota before it reaches later pages or higher-value operations.
- The same token is throttled from multiple machines, while a different account remains healthy.
How to diagnose it
Step 1
Capture status and policy headers from a safe request
Use a representative read and retain the response headers with a timestamp. Provider-specific quota headers can help, but interpret them using that provider’s documentation; similar names do not guarantee identical reset units or scopes.
curl -sS -D - -o /dev/null --max-time 15 https://api.example.com/itemsStep 2
Determine the identity the server is limiting
Compare the same account across workers, routes and schedules. Aggregate attempts by token identifier or tenant without logging raw secrets. If only traffic through a particular gateway is affected, inspect whether it is applying a source-IP policy.
Step 3
Calculate the offered load including retries
Record logical operations, actual HTTP attempts, queued work and completed successes separately. A growing attempts-per-success ratio reveals feedback that a request counter alone cannot explain. Include health probes, pagination and background polling.
Step 4
Check time handling with both valid formats
Test a numeric delay, a future HTTP date, a past date and an invalid value. If a trusted response Date is available it can help estimate clock skew; otherwise synchronise the client clock and avoid retrying early because of an incorrect local clock.
The fix
Share an admission limiter at the same scope as the server quota, or partition the allowance explicitly between workers. Maintain a bounded pending queue and backpressure the producer. A per-request sleep leaves every other request free to hit the same exhausted allowance.
Honour a valid Retry-After as a lower bound. Add random delay above it to spread the next wave; for missing or invalid values, use capped exponential backoff with jitter. If the required wait exceeds the operation’s deadline, defer or fail the operation instead of clamping the delay downward and retrying early.
Cap both attempts and total elapsed time. Give retries one owner so nested libraries do not multiply them. For mutations, retain an idempotency key across retries and follow the provider’s outcome semantics. Never infer from a transport or gateway response alone that an earlier attempt produced no side effect.
Reduce unnecessary demand: checkpoint pagination, avoid overlapping syncs, cache stable metadata and honour conditional requests where supported. When an allowance will not recover soon, persist the checkpoint and schedule continuation after reset rather than holding a worker and connection idle.
// JavaScript: null means the caller must use its bounded fallback policy.
function retryAfterMs(value, nowMs = Date.now()) {
if (value == null) return null;
const text = value.trim();
if (/^\d+$/.test(text)) {
const ms = Number(text) * 1000;
return Number.isSafeInteger(ms) ? ms : null;
}
// Date.parse accepts the standard HTTP-date emitted by conforming servers.
const when = Date.parse(text);
return Number.isFinite(when) ? Math.max(0, when - nowMs) : null;
}
// Compare this lower bound with the remaining deadline before scheduling.
// Add positive jitter; do not cap a valid server delay to an earlier retry.How to stop it coming back
- Test shared-token traffic from several workers, including the recovery burst after a cooldown.
- Expose rate-limit wait time and abandoned retries separately from dependency latency and successful throughput.
- Use bounded timers or durable scheduling for long delays rather than assuming every runtime accepts arbitrarily large timer values.
FAQ
Must a 429 include Retry-After?
No. It may include the header, but clients need a bounded fallback when it is missing. The response body or provider documentation may explain whether the limit is temporary or requires an account-level change.
Can I retry immediately with a different token?
Do not rotate identities to evade the service’s policy. Coordinate the authorised allowance, reduce demand or request an appropriate quota. Token rotation also makes incident accounting harder by hiding the real workload.
Does exponential backoff solve a sustained excess rate?
No. Backoff spreads recovery attempts, but a producer that continually offers more work than the allowance can service needs admission control or a lower workload. Otherwise the backlog grows without bound.
Related
Other errors engineers hit next to this one
- Permission denied when executing a script
- bad interpreter: No such file or directory with CRLF
- Argument list too long
- Too many open files
- Out of memory: Killed process
- No space left on device despite free disk space
- Text file busy during executable replacement
- set -e script continues after a failed pipeline