Node.js
Node.js — ECONNRESET: socket hang up on pooled keep-alive connections
Written and reviewed by Sahil Srivastav
Error: socket hang up
code: 'ECONNRESET'What this error actually means
A socket closed unexpectedly while the client was using it. With a reusable HTTP/1.1 connection, one particular race is common: the remote server or proxy retires an idle connection just as the client assigns it another request. The socket looked reusable when selected, but the peer’s close and the new write crossed in flight.
ECONNRESET is a transport observation, not proof of that race. A crashed upstream, an explicit local request destruction or a middlebox reset can produce similar output. The decisive clues are whether the request used a pooled socket, how long that socket had been idle, and whether a response had already begun.
A reset also leaves application outcome uncertain. The server may have received and committed a write before the connection disappeared. Retrying because no response arrived can duplicate a payment or job. Diagnosis of the transport and permission to repeat the operation are separate decisions.
Causes, most common first
- 1The peer closes idle connections before the client retires them. The effective peer may be a load balancer, sidecar or reverse proxy rather than the origin. Changing the origin’s timeout alone leaves the shorter middlebox timeout intact. There is always a small race near closure; compatible lifetimes reduce its frequency rather than proving it impossible.
- 2The upstream process is draining or restarting. Deployments can close pooled connections whose lifetime predates the deployment. Correlate resets with instance termination and readiness changes. An isolated reset after idleness suggests a different problem from every active connection dropping simultaneously.
- 3The caller aborts its own request. A deadline, request.destroy() or cancellation path can terminate the socket. Record the local abort reason alongside the network error; otherwise a correct cancellation is easily mistaken for an unhealthy server and triggers a pointless retry.
- 4Protocol or response handling prevents safe reuse. A mismatched HTTP/HTTPS endpoint, truncated body or unconsumed response changes the connection’s state. Confirm protocol and body completion first. Keep-alive tuning cannot repair a client that abandons response streams while assuming their sockets are free.
When you see it
- The first request after an idle interval fails, but the next fresh connection succeeds
- Failures concentrate on reused sockets and cluster near a proxy’s idle timeout
- Steady traffic works while sparse scheduled traffic sees intermittent resets
- Upstream logs may show no request, or may show a completed request without a delivered response
How to diagnose it
Step 1
Record whether Node reused the socket
For node:http or node:https requests, ClientRequest.reusedSocket identifies reuse. Log it together with error.code, elapsed time, method and whether a response began. This property belongs to the core HTTP client; fetch uses a different dispatcher and requires its own diagnostics.
Step 2
Observe assignment and closure in an isolated reproduction
Core HTTP debug logs can expose connection lifecycle, but may also expose request details. Use a local or sanitised staging reproduction, then send two requests separated by an idle interval around the suspected threshold.
NODE_DEBUG=http,net node reproduce.cjsStep 3
Map every connection boundary
List client-to-proxy and proxy-to-origin idle policies separately. Compare deployment configuration with running instances, because old workers may retain old settings. A request timeout, a TCP keepalive probe and an HTTP idle timeout are not interchangeable controls.
Step 4
Check the upstream outcome before replaying writes
Search server logs or the operation ledger by request ID or idempotency key. A missing response does not prove the server did nothing. For a state-changing call, establish an outcome lookup or deduplication contract before enabling automatic recovery.
The fix
Configure the client implementation to retire unused connections before the shortest known peer idle lifetime, with margin. Use options actually supported by the installed client; core Agent, third-party agents and Undici expose different controls. Do not copy an idle-timeout option from one implementation into another and assume it takes effect.
Keep pools bounded and long-lived. Creating an agent per request defeats reuse and increases handshakes and ephemeral-port churn. Disabling keep-alive briefly can help isolate the cause, but should not become the permanent solution without measuring the capacity cost.
Retry only a small, bounded number of times when the operation is safe to repeat and the caller’s deadline still permits it. A read may qualify under the API’s contract; a write needs a stable idempotency key and server-side deduplication, or an outcome check before resubmission.
For deployment-related failures, stop admitting new work before terminating instances and allow active requests to drain. Make cancellation propagate from callers so the service does not perform abandoned work while the retry creates a second attempt.
// Instrument core HTTP errors at the request boundary.
request.on("error", error => {
console.error({
code: error.code,
reusedSocket: request.reusedSocket,
method: request.method,
elapsedMs: Date.now() - startedAt,
});
// Retry policy also needs method semantics, remaining deadline,
// retry count and, for writes, a server-enforced idempotency key.
});How to stop it coming back
- Test idle gaps and rolling deployments, not only continuous load
- Track reset rate separately for new and reused connections where supported
- Document timeout ownership across every proxy hop
- Treat retry budgets and idempotency as part of the API contract
FAQ
Will TCP keepalive prevent this?
TCP keepalive probes detect some dead peers; they do not force an HTTP server or proxy to retain an idle connection. Align HTTP connection policies and preserve a safe retry strategy for the remaining closure race.
Can I retry every ECONNRESET once?
Only if repeating the operation is safe. A connection reset can occur after a successful server-side commit. Use the same idempotency key across attempts rather than generating a new key for the retry.
Why does this happen less under heavy traffic?
Frequent use can keep sockets away from the idle boundary. That pattern supports the stale-connection hypothesis, but confirm reuse and timing: traffic also changes proxy routing, instance selection and deployment exposure.
Related
Other errors engineers hit next to this one
- Replication lag: replica served stale data after a write
- command not found in a script that works interactively
- Permission denied when executing a script
- bad interpreter: No such file or directory with CRLF
- Argument list too long
- Too many open files
- Out of memory: Killed process
- No space left on device despite free disk space