Distributed systems
502 Bad Gateway — identify which upstream exchange failed
Written and reviewed by Sahil Srivastav
502 Bad GatewayWhat this error actually means
A gateway could not obtain a usable upstream response for this exchange. That gateway might be your CDN, a load balancer, nginx, or an application calling another service. The visible status identifies an intermediary failure; it does not identify the broken process. Start by finding which hop generated the response and which upstream address that hop selected.
The stage of failure determines the repair. A connection refused before any request bytes leave nginx suggests a listener, address or rollout problem. A connection accepted and then closed before response headers suggests an application crash or protocol mismatch. An oversized response header is a third case: the application answered, but the proxy could not accept the header block. These can all present as the same browser page.
An application can also deliberately return a 502. Compare the downstream status with the upstream status and error log before assuming nginx invented it. If retries are enabled, one client request can touch multiple upstreams; a final success can conceal a failing replica that will become an outage when the healthy capacity disappears.
Causes, most common first
- 1Traffic reaches an absent or incorrect listener. A service port changed, a process bound only to loopback inside a container, or readiness became true before the HTTP listener opened. A healthy node is not proof that the selected application address accepts connections.
- 2The upstream dies before sending headers. Worker termination, an uncaught error, memory exhaustion or a forced deployment shutdown ends the TCP exchange. Correlate process exit times with the selected upstream, rather than restarting every component and losing the evidence.
- 3The proxy speaks the wrong protocol. Sending plaintext HTTP to a TLS port, using the wrong SNI name, or treating a non-HTTP service as HTTP prevents a valid response. Reachability tests only prove that a port accepts a connection.
- 4The response cannot be parsed or buffered. Malformed headers and an unexpectedly large cookie or authentication header can fail before the body starts. The nginx error log distinguishes invalid headers from an oversized header buffer, which need different fixes.
When you see it
- Failures appear during deploys, while health checks against individual replicas still pass.
- One upstream address accounts for most errors; other addresses serve identical requests.
- The public endpoint fails while a direct request to the application succeeds.
- Login or large-cookie responses fail more often than small anonymous responses.
How to diagnose it
Step 1
Capture one request with its timing and headers
Replace the example host with your endpoint. Record the timestamp and request identifier from the response. A branded error body is a useful clue, but compare logs because intermediaries can replace each other’s bodies.
curl -sv --max-time 15 -o /dev/null -w '
status=%{http_code} connect=%{time_connect} first_byte=%{time_starttransfer} total=%{time_total}
' https://api.example.com/healthStep 2
Read the upstream-stage error on the proxy
On the nginx host, inspect the corresponding error window. Connection refusal, TLS handshake failure and premature closure provide substantially more information than the access-log 502. Restrict the window further on busy servers.
tail -n 200 /var/log/nginx/error.logStep 3
Probe the selected origin from the proxy network
Use the actual address from the error log and the same virtual-host name as production. This separates routing and listener failures from public-edge behaviour; a successful probe from your laptop does not test the proxy’s network path.
curl -sv --connect-timeout 2 --max-time 10 -H 'Host: api.example.com' http://127.0.0.1:8080/healthStep 4
Inspect the effective configuration
Check the upstream port, scheme, host forwarding and matching location together. nginx -T can expose credentials embedded in configuration, so inspect it locally. Confirm what is loaded rather than trusting an unused file in the repository.
nginx -TThe fix
Repair the failed stage. For refusal, correct the service address and bind interface, then gate readiness on the application’s ability to accept the required request. During shutdown, remove the instance from routing before closing its listener and allow in-flight requests a bounded drain period.
For a crash, fix the application exception or resource failure and confirm that the same input completes against the repaired instance. A proxy restart merely selects a different socket; it cannot make a worker finish an operation that kills it.
For a protocol mismatch, align the upstream scheme, TLS verification, SNI and virtual-host routing explicitly. For an oversized header, remove runaway cookies first; increase the specific header buffer only after measuring a legitimate bounded requirement.
Review retries before enabling them as mitigation. A failed response does not prove that a write did not commit. Limit retry attempts and elapsed time, and use durable idempotency for operations that can create duplicate effects. Inspect each upstream attempt, not just the final client status.
How to stop it coming back
- Log the selected upstream address, upstream status, connection time and header time alongside the request identifier.
- Exercise deployment draining and first-request readiness with ordinary traffic, not only a lightweight health route.
- Alert on per-replica upstream failures even when proxy retries turn the overall request into a success.
FAQ
Will raising proxy_read_timeout fix a 502?
Not a refused connection, invalid header or dead worker. Establish the failure stage first. A read timeout typically gives a different diagnostic message, and a larger wait can leave more requests consuming capacity while the real fault persists.
Why can a direct request work while nginx fails?
It may use a different network namespace, Host header, TLS name or replica. Reproduce the exact selected address from the proxy environment. A request to localhost on your laptop says nothing about localhost inside the proxy container.
Does a 502 mean the operation was rolled back?
No. The application may have committed and then lost the response. Resolve the outcome through an operation identifier or idempotency record before replaying a payment, reservation or other mutation.
Related
Other errors engineers hit next to this one
- awaitTermination never returns and the JVM will not exit
- InterruptedException caught and ignored — the task can no longer be cancelled
- Two unrelated components sharing a monitor via a boxed Integer or interned String
- Cache stampede — the same expensive value built many times concurrently
- Lock convoy — throughput collapses as threads are added, with no deadlock
- ReadWriteLock writer blocked indefinitely behind a stream of readers
- Worker loop never sees the stop flag and runs forever
- psycopg2.InterfaceError: connection already closed