Distributed systems

Upstream prematurely closed connection — the origin ended the exchange early

Written and reviewed by Sahil Srivastav

nginxUpstream lifecycleHTTP framing
upstream prematurely closed connection while reading response header from upstream

What this error actually means

nginx reached an upstream, then the connection ended before the expected HTTP exchange was complete. The quoted fragment is the response-header variant: nginx was still waiting for a complete header block. The phrase after while is diagnostic evidence, because losing an origin before headers and losing it halfway through a body have different client-visible outcomes.

Before downstream headers are sent, nginx may be able to produce a gateway error or retry under its configured rules. Once part of a response has been forwarded, it cannot replace the already-sent status with a clean 502 page. The client may receive status 200 followed by an incomplete body, so counting only HTTP status codes underestimates this failure.

Premature describes the protocol boundary, not necessarily a bug in nginx. A worker can be intentionally killed by its own timeout, terminated during deployment, or lost to an out-of-memory event. Those events all close a socket from the proxy’s perspective. The decisive evidence is the upstream process’s lifecycle at the same instant.

Causes, most common first

  1. 1A worker crashes or is killed while handling the request. An unhandled exception, native crash or memory kill prevents the worker from finishing headers or body. The request seen in the error log might be the victim of a process-wide resource problem rather than the request that originally caused it.
  2. 2The application server enforces a shorter worker timeout. A supervisor kills a worker before nginx’s read timeout expires. Increasing the proxy timer changes nothing because the origin no longer exists. Inspect the application server’s timeout log rather than inferring its budget from nginx configuration.
  3. 3Shutdown closes active connections without draining. A rollout stops the process while it still owns requests. Readiness withdrawal, load-balancer propagation and process termination have no coordinated grace period, so traffic reaches an instance that is already leaving.
  4. 4The service closes an unexpected protocol exchange. A wrong port, mismatched scheme, invalid request forwarding or application parser failure can make the peer close instead of producing HTTP headers. A listening socket by itself does not prove that the selected service speaks the expected protocol.

When you see it

  • nginx logs a premature closure and a selected upstream address for failed requests.
  • Failures line up with worker exits, restarts or a deployment’s shutdown window.
  • Large exports return partial data even though the access log records a successful status.
  • A specific input reliably terminates a worker while ordinary health checks pass.

How to diagnose it

Step 1

Keep the complete error line

Read the phase, upstream address, request path and timestamp together. A single extracted phrase loses the clues needed to match the process. Compare response-header failures with body failures instead of grouping them into one counter.

rg 'upstream prematurely closed connection' /var/log/nginx/error.log | tail -n 30

Step 2

Correlate the upstream worker’s lifecycle

For a systemd-managed example service, inspect exits and application exceptions around the same window. In containers use the equivalent previous-container logs and termination reason. A clean service restart can still kill an in-flight response.

journalctl -u api.service --since '15 minutes ago' --no-pager

Step 3

Check kernel evidence for a memory kill

Run on the host owning the upstream process. A container platform may expose the cgroup termination reason separately. Absence of this host log does not rule out a container memory limit or a supervisor-initiated kill.

journalctl -k --since '15 minutes ago' --no-pager | rg -i 'out of memory|killed process|oom'

Step 4

Compare one direct-origin exchange

From the proxy network, replay a safe failing read with the production Host header and a bounded client timer. If the origin also closes directly, focus on its lifecycle or protocol. Preserve the body and curl exit code when diagnosing a partial response.

curl -sv --max-time 30 -H 'Host: api.example.com' http://127.0.0.1:8080/export -o /tmp/export-check.bin

The fix

Fix the lifecycle event actually observed: bound memory for an unbounded export, repair the crashing path, or move a legitimately long operation into a job workflow. Avoid merely extending the worker timeout when the request holds unlimited memory or blocks indefinitely; that postpones termination while increasing resource occupancy.

For rollouts, stop admitting new work, allow routing changes to propagate, then drain in-flight requests within the platform’s termination grace. Give long streams an explicit reconnect or resume contract. A grace period that covers ordinary latency but not an unbounded stream needs a deliberate stream shutdown policy.

Correct protocol and host forwarding when the peer rejects the exchange. Reproduce with the exact scheme and endpoint from the error line; changing buffer sizes cannot repair a request sent to the wrong listener.

Treat partially delivered bodies as failed transfers. Clients should require valid framing and, where relevant, a checksum or complete-record marker before publishing a download. Retry only according to the operation’s safety contract; a lost response does not prove a mutation failed before committing.

How to stop it coming back

  • Record worker termination reasons and deployment identifiers alongside upstream request metrics.
  • Exercise a rolling restart during slow requests and streamed downloads, then verify complete bodies as well as status codes.
  • Alert on response truncation and upstream errors even when a retry or already-sent 200 hides the incident from status dashboards.

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Is this always a keep-alive problem?

No. Stale pooled sockets are one possible connection-lifecycle issue, but worker death, explicit application closure and forced shutdown produce similar symptoms. Establish whether failures follow idle reuse, active processing or a deployment before tuning keep-alive.

Can nginx retry after it has sent part of the response?

It cannot transparently replace a partially delivered response with a new complete response. Retry opportunities are constrained by response progress and request semantics. Clients need an explicit recovery strategy for truncated downloads or streams.

Why is there no application stack trace?

An external kill, segmentation fault or forced termination can stop the worker without a language-level exception log. Check the supervisor, kernel and container lifecycle evidence instead of concluding that no crash occurred.

Related

Other errors engineers hit next to this one

Full error and symptom index →