Distributed systems

Missing outbound timeout and cascading failure

Written and reviewed by Sahil Srivastav

Distributed systemsTimeoutsFailure propagation
requests.exceptions.ConnectionError: upstream request still pending after 120s

What this error actually means

An outbound call without a deadline converts a remote stall into a local resource leak. Each request holds a worker, connection, memory, and often a database transaction until the socket eventually closes. Once concurrency is consumed, healthy endpoints fail while their own dependencies remain healthy.

Timeouts must compose. The outbound deadline must be shorter than the inbound request deadline, which must leave time for cleanup and response. A proxy timeout alone is too late if application workers are already blocked.

The cascade is often mistaken for a capacity problem because adding workers briefly increases the runway. It also increases concurrent pressure on the stalled dependency.

Causes, most common first

  1. 1No connect or read timeout. A half-open route can occupy a caller indefinitely.
  2. 2Timeout longer than caller lifetime. The dependency deadline provides no useful cancellation.
  3. 3Connection held across remote work. A database slot is retained while an HTTP call waits.
  4. 4Retries multiply blocked calls. Each timeout creates more concurrent work.

When you see it

  • Active workers equal the concurrency limit
  • Pool pending counts rise while dependency throughput falls
  • Healthy endpoints fail together
  • Restarting instances temporarily restores service
  • Thread dumps show socket reads or pool waits

How to diagnose it

Step 1

Inspect blocked stacks

Sample workers and identify the call site holding capacity.

py-spy dump --pid <pid>
jstack <pid> | rg -n 'Socket|WAITING|pool'

Step 2

Measure pool queues

Compare active, idle, and pending resources at the failure time.

curl -s localhost:8080/metrics | rg 'threads|connections|pending'

Step 3

Trace dependency timing

Break latency into DNS, connect, first byte, and total duration.

The fix

Set connect, read, and total deadlines at the client boundary.

Release database and lock resources before waiting on unrelated services.

Propagate a deadline header and subtract elapsed time at every hop.

Cap concurrency to a value the dependency can sustain and shed excess work.

Retry only after cancellation is confirmed and the operation is safe.

deadline = monotonic() + 2.0
result = client.get(url, timeout=(0.2, max(0.01, deadline - monotonic())))

How to stop it coming back

  • Lint outbound calls without timeouts
  • Alert on pool pending before user errors
  • Inject blackholed and slow dependencies in tests
  • Track deadline remaining across hops
  • Keep remote calls outside database transactions

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Should the timeout equal the SLA?

No. It must leave time for downstream work, retries if allowed, cleanup, and the caller response.

Will more workers fix the cascade?

They delay exhaustion while increasing pressure. Fix the unbounded wait and bound concurrency.

Can TCP keepalive replace a read timeout?

No. Keepalive detects dead peers over a long interval; it does not bound an application that is alive but slow.

Related

Other errors engineers hit next to this one

Full error and symptom index →