Distributed systems

504 Gateway Timeout — find the timer that actually expired

Written and reviewed by Sahil Srivastav

HTTPTimeoutsLatency
504 Gateway Timeout

What this error actually means

A gateway stopped waiting for its upstream to complete the required exchange. A 504 is therefore a statement about one hop’s waiting policy, not proof that the whole backend is down. In a browser → CDN → load balancer → nginx → application chain, several timers can race; the first visible expiry often hides the slower work still running behind it.

Separate an overall deadline from an inactivity timeout. A total deadline limits elapsed time from the start of the operation. nginx proxy_read_timeout limits the interval between successive upstream reads; a response that keeps producing bytes can run longer than that value. Raising a read timeout cannot extend a separate CDN maximum request duration.

The more damaging failure is often after the 504: the backend continues a query or remote call, and the caller retries. Original and replacement requests now compete for the same pool. The status disappears only when work finishes or is cancelled, so a credible fix must address both the waiting budget and the lifetime of the work.

Causes, most common first

  1. 1Upstream work exceeds the endpoint’s budget. A changed query plan, lock wait, remote dependency or queue delay consumes the time available before headers arrive. The expensive stage can be several calls away from the proxy that reports the timeout.
  2. 2A connection cannot be established promptly. A dropped network path, full accept queue or unreachable origin consumes the connect budget. Distinguish this from connection refusal, which usually fails immediately rather than waiting until a timer expires.
  3. 3An inner timeout is longer than its caller’s deadline. The gateway gives up before the application’s database or HTTP timeout fires. Requests become abandoned work, and retries amplify the load because cancellation is not propagated to the resource actually doing the operation.
  4. 4Streaming pauses exceed an inactivity limit. An endpoint sends headers and then performs a long computation without writing. The connection can time out mid-body. Once headers have reached the client, the failure may appear as a truncated successful response rather than a new 504 status.

When you see it

  • Failures cluster around a repeatable elapsed duration, such as an example 30-second edge limit.
  • Application logs show successful completion after the gateway already returned a 504.
  • Small requests work while exports, lock waits or slow downstream operations time out.
  • Retry traffic increases after the first latency spike and delays recovery.

How to diagnose it

Step 1

Measure first-byte and total time

Run a safe representative request with a client limit longer than the suspected gateway threshold. Near-zero connection time and a long first-byte wait direct attention to upstream work; a client timeout shorter than the gateway masks the 504 entirely.

curl -sS --max-time 90 -o /dev/null -w 'status=%{http_code} connect=%{time_connect} first_byte=%{time_starttransfer} total=%{time_total}
' https://api.example.com/report/status

Step 2

Locate the timed-out stage in the proxy log

The error text distinguishes connecting to upstream, reading response headers and reading the response. Match the address and timestamp to application traces. Do not treat three different stages as one generic slow-server problem.

rg 'upstream timed out' /var/log/nginx/error.log | tail -n 30

Step 3

Inspect every configured timeout in the actual path

Review the loaded nginx configuration and the edge, client and application settings. Write down the meaning of each timer, including whether queueing and body consumption count. Equal numeric values with different semantics are not a coherent deadline policy.

nginx -T 2>&1 | rg 'proxy_(connect|read|send)_timeout|send_timeout'

Step 4

Follow a request beyond the visible failure

Use the request or trace identifier to find when downstream work really stops. If a database query continues for another minute, the missing cancellation or statement budget is evidence of the cascade, even if the original slow query also needs optimisation.

The fix

Build a budget from the outside inward. For an illustrative 10-second caller deadline, reserve time for queueing, response delivery and cleanup, then give downstream stages smaller bounds within the remaining time. Propagate an absolute deadline or recomputed remaining budget; do not restart a full timeout at every hop.

Remove the dominant wait: fix the query plan, shorten the transaction holding the lock, isolate the slow dependency, or reject excess work before it queues. Measure the latency distribution under concurrency after the change; one fast manual request does not demonstrate capacity.

Cancel request-owned work when the deadline expires and release pool leases on every outcome. Where cancellation cannot interrupt a database operation immediately, use a database-side execution bound and account for that work until it really terminates. Never return a still-busy connection as if it were safe to reuse.

For legitimate long jobs, use an accepted job plus a status resource or downloadable result. If an endpoint intentionally streams, define heartbeat and flush behaviour and configure each hop for that contract. A heartbeat should preserve an intended stream, not keep an unbounded stuck operation alive forever.

How to stop it coming back

  • Track queue time separately from dependency service time so a saturation incident is not mistaken for slow computation.
  • Test with a downstream that accepts connections but never responds, and assert that occupied slots return to baseline.
  • Use a single bounded retry budget within the caller deadline, with jitter and idempotency appropriate to the operation.

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Is a larger nginx timeout ever appropriate?

Yes, when the endpoint intentionally takes longer and every outer hop supports that contract. First establish resource bounds and cancellation. Extending a timer without those changes increases the number of requests that can wait simultaneously.

Can the server finish after the client receives 504?

Yes. Closing an HTTP exchange does not automatically stop database statements, background tasks or another service. Trace actual completion and design cancellation at those boundaries rather than assuming socket closure rolled back the operation.

Why do different clients report different errors?

Their budgets may expire at different hops. A short-lived client can disconnect first, leaving nginx to log a client abort; another waits long enough to receive the gateway’s 504. Compare the same operation and elapsed time.

Related

Other errors engineers hit next to this one

Full error and symptom index →