Python

Gunicorn WORKER TIMEOUT

Written and reviewed by Sahil Srivastav

PythonGunicornProduction incident
[2026-10-01 10:12:44 +0000] [1] [CRITICAL] WORKER TIMEOUT (pid:1234)

What this error actually means

Gunicorn’s arbiter did not receive a heartbeat from a worker within `timeout`. With sync workers, a request doing CPU or blocking I/O prevents the heartbeat; with async workers, a blocking call can stall the event loop. The arbiter kills the worker and in-flight requests are lost.

The timeout is measured by the server process, not by your endpoint’s business deadline. Raising it may stop restarts while making users wait longer and allowing one stuck dependency to consume capacity. The right fix starts by identifying what held the worker.

A worker timeout can also be a symptom of CPU throttling, stop-the-world garbage collection, or a deadlocked native extension.

Causes, most common first

  1. 1Blocking outbound call. A socket read or database call without a deadline occupies the worker.
  2. 2CPU-bound request. Large parsing, compression, or pure Python loops prevent heartbeats.
  3. 3Deadlock or native stall. A lock or extension holds execution indefinitely.
  4. 4Resource throttling. The container receives too little CPU to meet the heartbeat window.

When you see it

  • Workers restart at a regular timeout interval
  • Requests terminate together with the worker
  • CPU is saturated or near zero depending on the block
  • A dependency call has no timeout
  • The problem follows one endpoint or payload

How to diagnose it

Step 1

Capture a worker stack before the kill

Use Gunicorn’s worker-introspection signal or a debugger in staging to see the blocked frame.

kill -USR1 <gunicorn-master-pid>
# inspect the worker traceback in the Gunicorn log

Step 2

Check CPU and cgroup throttling

A timeout during CPU throttling is not an application latency change alone.

cat /sys/fs/cgroup/cpu.stat | rg 'nr_throttled|throttled_usec'
ps -o pid,pcpu,stat,wchan -p <worker>

Step 3

Trace network waits

Look for sockets with no response deadline and correlate their host with the endpoint.

strace -tt -p <worker> -e trace=network -f

The fix

Give every database and HTTP call a deadline shorter than the request and Gunicorn timeout.

Move CPU-heavy or long jobs to a queue; return a job identifier instead of holding a request worker.

Use an async worker class only when the application and all libraries are non-blocking; changing the class cannot make `requests` asynchronous.

Set worker count from measured CPU and memory, then reserve headroom for the arbiter and sidecars.

Use graceful shutdown and idempotent retries because a timeout can occur after a downstream side effect.

# gunicorn.conf.py
timeout = 60
graceful_timeout = 30
workers = 4
threads = 2

# application code still needs a shorter dependency deadline
response = requests.get(url, timeout=(2, 8))

How to stop it coming back

  • Alert on worker restarts and queue time, before timeout counts spike
  • Exercise slow dependencies and large payloads in load tests
  • Keep inbound timeout greater than dependency timeout plus cleanup margin
  • Profile CPU paths and event-loop lag
  • Make endpoint side effects idempotent

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Should I increase timeout?

Only when the measured legitimate request duration exceeds it and all dependencies are bounded. Otherwise you are hiding a blocked worker.

Why does async Gunicorn still timeout?

One synchronous library call can block the event loop just like a sync worker blocks its thread.

Does worker recycling prevent this?

It can contain leaks, but it cannot fix a request that exceeds the deadline.

Related

Other errors engineers hit next to this one

Full error and symptom index →