Distributed systems
Missing outbound timeout and cascading failure
Written and reviewed by Sahil Srivastav
requests.exceptions.ConnectionError: upstream request still pending after 120sWhat this error actually means
An outbound call without a deadline converts a remote stall into a local resource leak. Each request holds a worker, connection, memory, and often a database transaction until the socket eventually closes. Once concurrency is consumed, healthy endpoints fail while their own dependencies remain healthy.
Timeouts must compose. The outbound deadline must be shorter than the inbound request deadline, which must leave time for cleanup and response. A proxy timeout alone is too late if application workers are already blocked.
The cascade is often mistaken for a capacity problem because adding workers briefly increases the runway. It also increases concurrent pressure on the stalled dependency.
Causes, most common first
- 1No connect or read timeout. A half-open route can occupy a caller indefinitely.
- 2Timeout longer than caller lifetime. The dependency deadline provides no useful cancellation.
- 3Connection held across remote work. A database slot is retained while an HTTP call waits.
- 4Retries multiply blocked calls. Each timeout creates more concurrent work.
When you see it
- Active workers equal the concurrency limit
- Pool pending counts rise while dependency throughput falls
- Healthy endpoints fail together
- Restarting instances temporarily restores service
- Thread dumps show socket reads or pool waits
How to diagnose it
Step 1
Inspect blocked stacks
Sample workers and identify the call site holding capacity.
py-spy dump --pid <pid>
jstack <pid> | rg -n 'Socket|WAITING|pool'Step 2
Measure pool queues
Compare active, idle, and pending resources at the failure time.
curl -s localhost:8080/metrics | rg 'threads|connections|pending'Step 3
Trace dependency timing
Break latency into DNS, connect, first byte, and total duration.
The fix
Set connect, read, and total deadlines at the client boundary.
Release database and lock resources before waiting on unrelated services.
Propagate a deadline header and subtract elapsed time at every hop.
Cap concurrency to a value the dependency can sustain and shed excess work.
Retry only after cancellation is confirmed and the operation is safe.
deadline = monotonic() + 2.0
result = client.get(url, timeout=(0.2, max(0.01, deadline - monotonic())))How to stop it coming back
- Lint outbound calls without timeouts
- Alert on pool pending before user errors
- Inject blackholed and slow dependencies in tests
- Track deadline remaining across hops
- Keep remote calls outside database transactions
FAQ
Should the timeout equal the SLA?
No. It must leave time for downstream work, retries if allowed, cleanup, and the caller response.
Will more workers fix the cascade?
They delay exhaustion while increasing pressure. Fix the unbounded wait and bound concurrency.
Can TCP keepalive replace a read timeout?
No. Keepalive detects dead peers over a long interval; it does not bound an application that is alive but slow.
Related
Other errors engineers hit next to this one
- Connection reset by peer on a long-polling endpoint
- Redis OOM command not allowed above maxmemory
- MISCONF Redis is configured to save RDB snapshots
- READONLY You can’t write against a read only replica
- LOADING Redis is loading the dataset in memory
- CROSSSLOT Keys in request don’t hash to the same slot
- MOVED and ASK replies from Redis Cluster
- Redis clients stall during KEYS on a large keyspace