Distributed systems
Request timeouts cascade when the underlying work keeps its pool slot
Written and reviewed by Sahil Srivastav
upstream timed outWhat this error actually means
A caller’s timeout stops that caller waiting; it does not necessarily stop the operation that borrowed a database connection, worker thread or outbound HTTP slot. The visible nginx fragment upstream timed out is one possible symptom of this cascade. The actual client and pool exception text varies by stack, so diagnose the lifetime relationship rather than matching one exception class.
Imagine a service with 20 database slots and a two-second caller deadline. If statements continue for 30 seconds after callers leave, those slots remain occupied while replacement requests accumulate. This example needs no memory leak: resources can eventually return and the service can still be unavailable because occupancy far exceeds the useful request lifetime.
Retries then make the feedback loop worse. New attempts enter the same queue behind abandoned work, exceed their own budgets and become more abandoned work. Increasing pool size lets more operations reach the already-slow dependency and can move the bottleneck into the database, turning a local problem into a shared outage.
Causes, most common first
- 1The timeout races a future without cancelling its work. A wrapper returns an error after a timer wins, but the original promise, thread or driver call continues. The HTTP handler looks finished while its downstream resource remains borrowed, often outside the visibility of request metrics.
- 2Pool acquisition and queueing are outside the deadline. The application gives each stage a fresh full timeout after waiting for a slot. A request that already spent its useful budget in a queue still begins an expensive operation whose caller will shortly leave.
- 3A connection is held across unrelated remote work. An open transaction waits for an HTTP service or a user-level retry delay. A slowdown in that remote service monopolises database slots even though the database itself has little useful work to perform.
- 4Retries and cleanup have no shared ownership. SDK, handler and gateway retries multiply attempts, while exceptional branches skip response-body closure or transaction cleanup. Genuine leaked resources can coexist with slow-but-eventually-released work, so verify both patterns.
When you see it
- Client timeouts occur first, then pool-acquisition timeouts spread to unrelated endpoints.
- In-flight database or HTTP operations remain high after client traffic stops.
- The system recovers much later than the caller deadline would suggest.
- Actual dependency attempts grow faster than completed logical requests after retries start.
How to diagnose it
Step 1
Compare request completion with resource release
Record request start, deadline, cancellation, downstream start/end and lease return on one trace. The critical interval is between the caller leaving and the underlying operation terminating. A cancellation log alone does not prove that interval is zero.
Step 2
Ask PostgreSQL what active and abandoned sessions hold
For a PostgreSQL-backed service, inspect states and wait events during the incident. Query age shows live statement duration; idle-in-transaction sessions point to a different transaction-lifecycle fault. Restrict to the application identity when several services share the database.
SELECT pid, state, wait_event_type, wait_event,
now() - query_start AS query_age,
now() - xact_start AS transaction_age
FROM pg_stat_activity
WHERE datname = current_database() AND pid <> pg_backend_pid()
ORDER BY query_start NULLS LAST;Step 3
Measure pending acquisitions and active leases together
Export pending waiters, occupied slots, acquisition duration and release count from the relevant pool. A flat maximum occupancy with growing pending requests indicates saturation; leases that never return after all operations terminate indicate a lifecycle leak.
Step 4
Inject a dependency stall in a controlled test
Make a stub accept a request and stop responding. Send only a bounded number of calls, let their deadlines expire and assert that pending queues and resource occupancy recover. Then repeat with a client disconnect and with one retry enabled to expose amplification.
The fix
Carry one operation deadline through queueing, pool acquisition and downstream work. Before each stage, compute remaining time and reject already-expired work. Reserve a cleanup margin so a dependency can finish or cancel before the outer caller’s limit; a sequence of fresh full timeouts defeats the total budget.
Use cancellation mechanisms that reach the actual I/O or statement, and release the lease in a finally-style scope after the operation settles. With an HTTP client, consume or close response bodies according to that client’s pooling contract. With a database, cancel or bound the statement and roll back a failed transaction before reusing its connection.
Do not hold database transactions while waiting for unrelated network operations. Shorten the lease scope, move remote work outside it and use explicit workflow or outbox semantics where both systems must coordinate. Simply returning the connection early while its statement still executes risks concurrent use of one session.
Bound queues and concurrency per dependency, and reject overload before requests become too old to complete. Give retries one owner, a small attempt budget and a share of the original deadline. Where a dependency is failing persistently, a circuit breaker can stop new attempts while recovery is measured.
How to stop it coming back
- Track useful completions per borrowed-resource second, so abandoned work is visible even when release eventually succeeds.
- Require timeout tests to assert cancellation and resource recovery, not merely that the caller receives an error promptly.
- Capacity-plan across all application instances rather than multiplying a generous local pool by autoscaling count.
FAQ
Is Promise.race with a timer enough?
It can bound the caller’s wait but does not cancel the losing operation by itself. Pass a supported cancellation signal to the underlying client and verify that the operation and its resources actually terminate on timeout.
Should I kill every long query to recover?
First identify the affected application and the operation’s transaction semantics. Targeted cancellation can relieve an incident, but indiscriminate termination can disrupt unrelated work and trigger more retries. Fix admission and timeout ownership so the backlog does not immediately reform.
How is this different from a connection leak?
A leak retains a resource after the operation should have released it, potentially forever. This cascade can happen even when every resource eventually returns: it simply stays busy far beyond the caller’s useful deadline. Observe recovery after work truly stops to separate them.
Related
Other errors engineers hit next to this one
- Undici / fetch connections remain occupied
- ERR_UNHANDLED_REJECTION in a worker thread
- command not found in a script that works interactively
- Permission denied when executing a script
- bad interpreter: No such file or directory with CRLF
- Argument list too long
- Too many open files
- Out of memory: Killed process