Distributed systems
Connection reset by peer during long polling
Written and reviewed by Sahil Srivastav
Connection reset by peerWhat this error actually means
The local socket reports that its peer reset the TCP connection, so the current exchange cannot continue. The peer may be the origin, a reverse proxy or a network device on the path; the text does not identify which application made the decision. Operating-system and client wrappers add their own context, such as a receive failure or ECONNRESET.
A long poll intentionally holds an active HTTP request until an event arrives or a hold timer expires. During a quiet period, no application bytes may cross the connection. An intermediary with a shorter inactivity limit can close that exchange even though the application is behaving exactly as designed.
Do not confuse this with an idle keep-alive connection between requests. Here a request is outstanding and waiting for its response. Client read limits, proxy upstream inactivity limits and load-balancer active-request policies matter; extending only the pool’s between-request idle lifetime may have no effect.
Causes, most common first
- 1The poll hold time exceeds a hop’s inactivity budget. The application waits longer for an event than the client, gateway or firewall permits silence. When events arrive frequently the mismatch remains hidden because each response completes before the shorter timer expires.
- 2Shutdown aborts active polls. Long-lived requests outlast the ordinary deployment grace period. The process or intermediary closes them abruptly, and clients immediately retry together. A rollout becomes a repeatable capacity event rather than a small transient interruption.
- 3Cancellation leaves stale waiters registered. The client disconnects, but the server retains its subscription callback or occupied worker. Reconnecting creates another waiter for the same consumer; enough abandoned waiters exhaust threads, descriptors or memory and provoke further resets.
- 4A healthy empty poll is misclassified as failure. The protocol lacks an explicit no-event completion and retry delay. Clients interpret routine hold expiry as a transport problem, overlap new polls, or discard their event cursor and create gaps or duplicates during recovery.
When you see it
- Quiet subscriptions fail near the same elapsed time, while busy subscriptions complete normally.
- Direct-origin polling works but the public route through a gateway resets.
- Deployments disconnect large groups of pollers, followed by a reconnection burst.
- Server-side waiter counts grow when clients reconnect without cancelling earlier polls.
How to diagnose it
Step 1
Measure a deliberately quiet poll
Use a staging channel with no events and a client limit longer than the intended hold. This example requests a 20-second wait; use the actual API’s parameter. Compare first-byte and total time and preserve curl’s nonzero exit result on a reset.
curl -sv --max-time 35 -o /dev/null -w '
status=%{http_code} first_byte=%{time_starttransfer} total=%{time_total}
' 'https://staging.example.com/events?wait=20'Step 2
Compare public and direct-origin paths
Run the same quiet poll from the proxy network against the selected origin, preserving authentication and virtual-host routing. If only the public path fails at a stable threshold, inventory the timers on each additional intermediary.
Step 3
Separate an explicit HTTP timeout from a TCP reset
Inspect gateway logs and client output. A 504 response, a clean empty response and a reset are different outcomes. The gateway may produce a 504 on its downstream leg while closing its upstream socket, so correlate both ends of one request.
Step 4
Observe resets at the relevant network leg
On an authorised Linux staging host, this packet summary selects TCP reset packets for port 8080. Choose the real upstream port and interface. It can show the sender address, but NAT and intermediaries mean that address alone may not identify the responsible process.
sudo tcpdump -nn -i any 'tcp port 8080 and (tcp[tcpflags] & tcp-rst != 0)'Step 5
Count live waiters after client disconnects
Open a small known number of polls, cancel them, and verify listener registrations, worker slots and subscription state return to baseline. A reconnect test that checks only successful events misses retained abandoned polls.
The fix
Make the server’s poll hold interval shorter than every supported intermediary inactivity limit, with time left for scheduling and response delivery. Return a defined empty result on normal hold expiry. For an illustrative shortest path limit of 30 seconds, a 20-second hold can leave margin; measure the actual path instead of relying on those example values.
Set the client deadline slightly beyond the intended hold plus transport margin and avoid overlapping polls for one logical subscription. An empty result should continue from the same cursor; a reset should reconnect with bounded backoff and jitter while retaining the last acknowledged position.
Unregister request-scoped event listeners and release resources on completion, cancellation and exceptions. If a framework uses an asynchronous waiter, test that its cancellation reaches the subscription registry. Parking a dedicated thread per quiet client can exhaust a bounded executor even when CPU is low.
For deployments, stop accepting new polls and complete existing ones with an ordinary retryable protocol outcome before termination. Stagger reconnects and ensure the client can resume from a durable cursor. If continuous streaming is the intended contract, use an explicit streaming protocol with heartbeat and buffering rules rather than casually injecting bytes into a JSON long-poll response.
How to stop it coming back
- Test the longest quiet interval through the real gateway chain, not only the low-latency event-delivery case.
- Load-test cancellation and reconnect storms while measuring active waiters, open sockets and listener counts.
- Define cursor retention, duplicate handling and replay limits so reconnecting after a reset preserves event correctness.
FAQ
Will TCP keepalive stop the long poll from timing out?
Not necessarily. TCP probes can help detect dead transport peers or influence some network devices, but an HTTP proxy may still enforce its own application-read inactivity or total-request limit. Align the poll contract with those timers explicitly.
Should I send whitespace as a heartbeat?
Only if the response format and client parser intentionally support streaming and each intermediary flushes it. Bytes sent after headers turn the exchange into an active body stream; they are not a universal fix for a long-poll endpoint expecting one complete JSON result.
Can a reset make a consumer miss an event?
Yes, if the server advances delivery state before the client durably acknowledges the event and reconnection starts from the wrong position. Use a stable cursor or acknowledgement protocol and tolerate duplicates instead of equating a successful socket write with durable consumption.
Related
Other errors engineers hit next to this one
- psycopg2.InterfaceError: connection already closed
- RecursionError: maximum recursion depth exceeded
- MemoryError: unable to allocate array
- QueuePool limit of size 5 overflow 10 reached, connection timed out
- DetachedInstanceError: instance is not bound to a Session
- RuntimeError: Event loop is closed
- Task was destroyed but it is pending!
- Executing <Handle ...> took 2.418 seconds (blocked event loop)