Distributed systems
Intermittent 502 after idle — test for stale connection reuse
Written and reviewed by Sahil Srivastav
502 Bad GatewayWhat this error actually means
A proxy can retain an upstream connection for reuse after the previous response finishes. If the origin or an intermediate device retires that idle connection earlier, the proxy can race with the close while assigning the next request. The connection looked reusable when selected, but the exchange fails before a usable response arrives.
This is a hypothesis to test, not a diagnosis implied by every intermittent 502. The characteristic pattern is a failure on the first request after an idle gap, followed by a successful request on a fresh connection. Sustained traffic may conceal the issue because sockets never remain idle long enough to cross the mismatched lifetime.
Keep-alive settings are directional. A client-facing nginx keepalive_timeout controls browser-to-nginx idle connections, whereas an upstream keepalive_timeout controls nginx-to-origin cached connections. Neither is the same as proxy_read_timeout during an active response. Editing the wrong timer leaves the failing socket’s lifetime unchanged.
Causes, most common first
- 1The pool keeps idle sockets longer than the peer. The origin’s idle close can arrive near the instant the proxy reuses the socket. Well-behaved implementations discard observed closed sockets, but cannot eliminate every race between selection, close delivery and the next write.
- 2An intermediate device expires idle state earlier. A firewall, NAT or load balancer sits between the pool and origin. Its inactivity policy can dominate both application settings, so apparently aligned endpoints still fail after a repeatable idle gap.
- 3The wrong direction’s timeout was changed. An operator extends browser-facing keep-alive or active-response read timeouts while the upstream cache still retains sockets beyond the origin’s policy. Similar directive names make this an easy configuration-review mistake.
- 4Connection retirement or deployment races are mistaken for idle expiry. Maximum connection age, maximum requests per connection or a process rollout can close a socket independently of idle time. Compare failures with socket age and deployment events before asserting that only an idle timeout is involved.
When you see it
- The first request after a quiet period fails; an immediate second request succeeds.
- Failures increase when traffic becomes sparse rather than during peak throughput.
- A direct-origin request opening a new connection works consistently.
- The problem appears after changing origin server defaults, a load balancer or connection-pool settings.
How to diagnose it
Step 1
Inspect upstream and client-facing contexts separately
Read the enclosing blocks in the effective configuration. Record the upstream idle timeout, pool size and origin keep-alive settings. The upstream keepalive count limits cached idle connections per worker; it is not a global cap on all connections.
nginx -TStep 2
Vary the idle gap against a safe staging endpoint
Send two requests through the same proxy, varying the pause near the suspected origin idle limit. Each curl process creates a fresh client-facing connection, while nginx can still reuse its own upstream pool. Repetition may be needed because several workers or pooled sockets exist.
curl -sS --max-time 10 https://staging.example.com/health
sleep 6
curl -sv --max-time 10 https://staging.example.com/healthStep 3
Correlate reuse, closure and selected upstream
Use proxy connection timing and error logs, then a controlled packet capture if necessary. A new connection has a handshake; a reused connection does not. A FIN or reset near the failing reuse supports the hypothesis, but identify which network leg emitted it.
Step 4
Compare with reuse disabled in staging
Temporarily remove the upstream idle cache in the test configuration and replay the same idle-gap workload. Disappearance of failures is useful evidence, not a reason to permanently discard pooling without assessing the connection and TLS costs.
The fix
Set the pool’s idle retention shorter than the earliest known peer or middlebox idle expiry, leaving a margin for scheduling and timing variation. For an illustrative origin idle policy of 30 seconds, an upstream cache timeout of 20 seconds makes the pool retire sockets first. Use measured policies rather than copying these numbers as universal defaults.
Make the HTTP version and Connection header behaviour explicit for upstream persistence, and keep origin settings consistent across replicas. Rolling out only half the servers creates a mixed policy that produces intermittent failures tied to selected addresses.
Keep a bounded retry for eligible operations within the original deadline, because network-close races can still happen. Do not enable retries of arbitrary non-idempotent writes merely to hide 502s. The peer may have read and committed a request before the response was lost.
If infrastructure has an unavoidable short idle expiry, shorten the pool retention or redesign the network path. TCP keepalive probes and HTTP keep-alive are different mechanisms; enabling socket probes does not automatically change an HTTP server’s idle request policy.
# Example nginx http-context configuration. Origin idle timeout: 30s.
upstream api_keepalive {
server 127.0.0.1:8080;
keepalive 32;
keepalive_timeout 20s;
}
server {
listen 8081;
location / {
proxy_pass http://api_keepalive;
proxy_http_version 1.1;
proxy_set_header Connection "";
}
}
# Run nginx -t and verify the origin/middlebox policy before adopting values.How to stop it coming back
- Document idle timeout, maximum connection age and request-count limits for each network leg.
- Include sparse traffic with idle gaps in connection-pool regression tests; continuous load misses this failure shape.
- Measure new connection rate alongside errors so a workaround does not silently create connection churn or port pressure.
FAQ
Does a Connection: close request from curl disable nginx upstream reuse?
Not necessarily. That header describes the client-facing exchange and nginx controls forwarding and its upstream connection pool separately. Change the relevant upstream test configuration when isolating reuse.
Should the proxy or the origin close idle sockets first?
For a client pool reusing origin connections, it is generally preferable for the pool to retire its cached socket before the peer’s idle expiry. That reduces the chance of assigning new work to a socket the peer is closing.
Will increasing proxy_read_timeout help?
No. An already closed socket does not become valid by waiting longer. Read timeouts bound inactivity during an active exchange; this incident concerns whether an idle connection is safe to reuse at all.
Related
Other errors engineers hit next to this one
- Connection reset by peer on a long-polling endpoint
- Redis OOM command not allowed above maxmemory
- MISCONF Redis is configured to save RDB snapshots
- READONLY You can’t write against a read only replica
- LOADING Redis is loading the dataset in memory
- CROSSSLOT Keys in request don’t hash to the same slot
- MOVED and ASK replies from Redis Cluster
- Redis clients stall during KEYS on a large keyspace