Python
Blocking calls inside async def starving the event loop
Written and reviewed by Sahil Srivastav
WARNING:asyncio:Executing <Handle <TaskWakeupMethWrapper object at 0x7f1c> created at /app/api.py:88> took 2.418 seconds
WARNING:asyncio:Executing <TimerHandle when=41233.9 Server._on_timeout() at /app/api.py:214> took 1.902 secondsWhat this error actually means
An asyncio application runs all your coroutines on one thread. Concurrency comes entirely from coroutines yielding at `await` points; between two `await`s, your code owns the thread exclusively. A synchronous call that takes 200 milliseconds therefore does not cost 200 milliseconds of latency to one request — it adds 200 milliseconds to every request currently in flight, plus everything queued behind them, plus the health check, plus the heartbeat that keeps your Kafka or database connection registered.
That multiplication is why a blocked loop presents as a whole-service problem with no single slow endpoint. Percentiles rise together, timeouts appear in unrelated code, and the traces show time spent waiting rather than working. The endpoint doing the blocking often looks fine in its own metrics.
This is also the most easily introduced async bug, because the blocking call is usually correct code used in the wrong context: `requests.get` instead of an async client, `psycopg2` instead of `asyncpg`, `time.sleep` instead of `asyncio.sleep`, a `bcrypt` hash, a 30 MB `json.loads`, a Pillow resize, a `pandas` group-by. None of them raise. None of them look like a concurrency mistake. The only signal is lag.
asyncio will tell you, if you ask. In debug mode the loop times every callback and logs `Executing <Handle ...> took N seconds` for anything exceeding `slow_callback_duration` (0.1s by default). That log line names the source location — it is the cheapest possible diagnostic for this entire class of problem.
Causes, most common first
- 1A synchronous network client inside a coroutine. The most common version. `requests`, `urllib`, a sync database driver, a sync Redis client, `boto3`, or any SDK without an async interface. The loop is frozen for the entire round trip, which under a slow dependency is seconds, not milliseconds.
- 2CPU-bound work on the loop thread. Password hashing, compression, image processing, large JSON serialisation, template rendering over big datasets, `pandas` or `numpy` operations. Legitimate work in the wrong place. Note that `await` does not help here — there is nothing to wait for, only work to do.
- 3time.sleep, blocking locks, or blocking queue operations. `time.sleep`, `threading.Lock.acquire`, `queue.Queue.get`, `subprocess.run` and `os.system` all park the thread. `asyncio.sleep`, `asyncio.Lock`, `asyncio.Queue` and `asyncio.create_subprocess_exec` are the non-blocking equivalents.
- 4Filesystem I/O treated as free. Reading configuration, writing logs synchronously, or reading an uploaded file is fine at small sizes and blocks measurably on network filesystems, large files, or a busy disk. There is no non-blocking file I/O in CPython’s default loop — it has to go to a thread.
- 5Implicit IO from a lazy ORM attribute. With async SQLAlchemy, touching an unloaded relationship raises `MissingGreenlet` rather than blocking. With a sync engine used from async code, the same access silently blocks the loop instead. The second failure is much harder to notice, and it is the one that degrades production.
When you see it
- p50 latency across every endpoint moves together, including endpoints that do almost nothing
- One CPU core pinned at 100% while the others idle, and throughput far below what the hardware should give
- Health checks and readiness probes time out intermittently, causing restarts under load
- Latency scales with concurrency rather than with the work each request does
- Adding more workers helps roughly linearly, which hides the bug and multiplies its cost
How to diagnose it
Step 1
Turn on debug mode and read the source locations
This is the decisive step and takes one environment variable. The loop reports every callback slower than `slow_callback_duration` with the file and line that created it. Lower the threshold to catch smaller offenders.
PYTHONASYNCIODEBUG=1 python -X dev app.py
# or, in code:
loop = asyncio.get_running_loop()
loop.set_debug(True)
loop.slow_callback_duration = 0.05Step 2
Measure loop lag continuously in production
Schedule a timer that should fire every interval and record how late it actually fires. That delta is event-loop lag, and it is the single most useful metric an asyncio service can export — it stays near zero on a healthy loop and tracks blocking directly.
async def lag_monitor(interval=0.25):
while True:
start = time.perf_counter()
await asyncio.sleep(interval)
LAG.observe(time.perf_counter() - start - interval)Step 3
Sample the stack of the loop thread while it is slow
A sampling profiler that can attach to a running process shows which frame the single thread is actually executing. Because everything runs on one thread, whatever is on top during a lag spike is the blocking call.
py-spy dump --pid <pid>
py-spy top --pid <pid>Step 4
Check whether latency scales with concurrency
Send one request and measure; send fifty concurrently and measure again. A non-blocking handler barely changes. A blocking one multiplies by roughly the concurrency, because the requests execute serially in effect.
hey -n 200 -c 1 http://localhost:8000/report
hey -n 200 -c 50 http://localhost:8000/reportThe fix
For I/O, use a client that yields. `httpx.AsyncClient` or `aiohttp` instead of `requests`, `asyncpg` or SQLAlchemy’s async engine instead of `psycopg2`, `redis.asyncio` instead of the sync client, `aioboto3` where the AWS SDK is on a hot path. This is the real fix: the loop stays responsive because the call actually suspends.
For work that cannot yield, move it off the loop thread with `await asyncio.to_thread(fn, *args)` — a thread pool suits blocking I/O and C extensions that release the GIL. For genuinely CPU-bound Python, a thread is not enough because of the GIL; use `loop.run_in_executor(ProcessPoolExecutor(), ...)` and accept the serialisation cost, or hand the work to a task queue.
Bound the offload. An unbounded thread pool converts a loop-blocking bug into a thread-explosion bug, and `asyncio.to_thread` uses the default executor whose size is limited but shared with everything else. Give heavy work its own executor with an explicit maximum and a queue, so overload produces backpressure rather than resource exhaustion.
Use the framework’s escape hatch instead of fighting it. In FastAPI, a handler declared `def` rather than `async def` runs in a threadpool automatically — so a legacy endpoint full of sync calls is safer as a plain `def`. Declaring it `async def` and leaving the sync calls inside is the single most common way this bug enters a FastAPI codebase.
Set timeouts everywhere. On a blocked loop, a missing timeout on one dependency becomes a service-wide stall rather than one slow request, so timeouts are load-bearing for isolation and not merely for tidiness.
# Freezes every concurrent request for the duration of the call
@app.get("/report")
async def report(user_id: int):
resp = requests.get(f"{BILLING}/usage/{user_id}", timeout=10) # blocks the loop
digest = bcrypt.hashpw(resp.content, bcrypt.gensalt()) # blocks the loop
return {"digest": digest.decode()}
# Async client for I/O; CPU work moved to a bounded executor
CPU = ProcessPoolExecutor(max_workers=2)
@app.get("/report")
async def report(user_id: int):
resp = await http.get(f"{BILLING}/usage/{user_id}", timeout=10.0)
loop = asyncio.get_running_loop()
digest = await loop.run_in_executor(CPU, hash_bytes, resp.content)
return {"digest": digest}How to stop it coming back
- Export event-loop lag as a first-class metric and alarm on it; it detects this class of regression within one deploy
- Run tests and local development with `-X dev` and `PYTHONASYNCIODEBUG=1` so slow callbacks are visible before review
- Forbid sync clients in async modules with a lint rule — `requests`, `time.sleep`, `psycopg2` imported inside async code should fail CI
- Declare handlers `def` when they contain sync work and `async def` only when they are genuinely non-blocking; do not mix inside one function
- Load-test at realistic concurrency, not one request at a time; serial latency looks perfect on a blocked loop
FAQ
Does adding await in front of a synchronous call help?
No, and it will not even compile unless the call returns an awaitable. `await` marks a suspension point; it cannot create one. A synchronous function runs to completion on the loop thread regardless of what keyword precedes it.
What is a healthy event-loop lag?
Single-digit milliseconds at p99 for a request-serving service, with p50 effectively zero. Sustained tens of milliseconds means something is regularly hogging the thread; hundreds means your concurrency is nominal. Track the number rather than a threshold someone quoted, because it is comparable across deploys.
Is asyncio.to_thread enough for CPU-bound work?
Only if the work releases the GIL — much of `numpy`, compression and C-extension I/O does. Pure Python computation in a thread still contends for the GIL, so the loop thread is starved almost as badly. That case needs a process pool or an external worker.
Why did this only appear under load?
Because the cost is proportional to how many requests are waiting. At concurrency one, a 200 ms block is a 200 ms request. At concurrency fifty, it is up to ten seconds of queueing distributed across the service, which is when timeouts and failing health checks start.
Related
Other errors engineers hit next to this one
- 429 Too Many Requests and Retry-After
- nginx 499 client closed request
- upstream prematurely closed connection
- Intermittent 502 after an idle keep-alive connection
- SSL certificate problem: unable to get local issuer certificate
- ERR_INCOMPLETE_CHUNKED_ENCODING
- Request timeouts cascading into pool exhaustion
- Connection reset by peer on a long-polling endpoint