Node.js
Node.js — Event loop blocked by synchronous work
Written and reviewed by Sahil Srivastav
What this error actually means
JavaScript callbacks in one Node event loop run until they return or yield through an actual asynchronous boundary. A callback that parses a huge document, runs a costly regular expression or performs synchronous filesystem work prevents unrelated callbacks from running. Ready sockets and expired timers must wait behind that work.
There is no standard runtime exception for a blocked event loop, so this symptom has no fabricated error message. The service may remain alive and report no application errors while latency rises across otherwise unrelated endpoints. A timer is not a real-time deadline: its callback cannot execute while JavaScript is monopolising the loop.
High CPU is useful evidence but not a complete definition. A synchronous child process can block the loop while the Node process itself uses little CPU. Conversely, asynchronous crypto or filesystem operations can queue in the worker pool while the main loop remains responsive. Measure loop delay and operation latency separately before changing thread counts.
Causes, most common first
- 1Input-dependent CPU work runs in a request callback. Large JSON parsing, serialisation, sorting or nested loops scale with payload size. A route that is fast for a small fixture can block every request for the largest tenant. Async function syntax changes the return type; it does not move computation to another thread.
- 2A synchronous API is used after startup. readFileSync, synchronous compression or crypto, and child_process synchronous APIs block the calling thread. Startup configuration reads may be acceptable before serving traffic; the same operation inside a frequently executed route has a different latency cost.
- 3Microtasks or nextTick callbacks never let the loop advance. A self-replenishing promise or nextTick chain can starve timers and I/O. Awaiting an already resolved promise repeatedly is not the same as yielding to the next event-loop phase. A batch loop must create a real scheduling boundary when it promises fairness.
- 4Runtime or host pauses dominate the measurement. Garbage collection, CPU throttling and host scheduling can delay callbacks even when application code has no obvious long loop. Correlate memory pressure and container CPU limits with profiles. Moving code to a worker will not create CPU capacity that the container does not have.
When you see it
- Health checks and cheap routes stall when one expensive endpoint runs
- Timers fire late and request timeouts arrive in clusters after a pause
- One core is busy even when host-wide CPU usage looks moderate
- Database query execution is fast but application response latency is high
How to diagnose it
Step 1
Install a delay histogram in the affected process
monitorEventLoopDelay returns nanoseconds. Divide by one million for milliseconds and reset between reporting windows. Include the maximum as well as a percentile: a single rare stall can matter even when most samples are normal. The sampling resolution contributes to the baseline.
Step 2
Capture a CPU profile during a controlled reproduction
Run the application with representative input sizes and then stop it normally to write the profile. Inspect stacks for parsing, sorting, regex work and user loops. An idle CPU profile during high lag points towards blocking system calls, scheduling or external pauses.
node --cpu-prof server.jsStep 3
Find synchronous calls in the hot path
This search is a starting point, not proof that every match is harmful. Follow route imports and distinguish startup-only work from per-request execution. Include application-specific CPU transformations that have no Sync suffix.
rg -n "readFileSync|writeFileSync|execSync|spawnSync|pbkdf2Sync|JSON\.(parse|stringify)" srcStep 4
Correlate lag with resource and payload metrics
Compare loop delay, single-process CPU, memory pressure, active requests and input bytes on the same timeline. A steep latency increase above a payload threshold suggests an algorithm or parsing bound; a plateau at a CPU quota suggests admission or capacity limits.
The fix
Replace request-path synchronous I/O with asynchronous APIs and await them through the request boundary. This helps only for APIs that actually perform work asynchronously; wrapping a synchronous function in Promise.resolve leaves it on the same loop.
Move substantial CPU tasks to a bounded worker-thread pool or an external job system. Bound queued tasks as well as active workers. Account for serialisation and transfer costs, and avoid sending a huge object graph to workers merely to move the memory problem elsewhere.
For work that can be partitioned, process a bounded chunk and yield with setImmediate before continuing. Check cancellation between chunks. Select chunk size by latency measurement; an arbitrary item count can still contain one exceptionally expensive item.
Enforce input and concurrency limits before expensive parsing or transformations. Where profiling shows an algorithmic problem, reduce the work itself with indexing, precomputation or a better algorithm. Adding replicas can increase throughput but does not reduce the pause one oversized request causes on its chosen instance.
import { monitorEventLoopDelay } from "node:perf_hooks";
const lag = monitorEventLoopDelay({ resolution: 20 });
lag.enable();
setInterval(() => {
console.log({
eventLoopP99Ms: lag.percentile(99) / 1e6,
eventLoopMaxMs: lag.max / 1e6,
});
lag.reset();
}, 10_000).unref();
// Partition suitable CPU work; each individual item must also be bounded.
for (let i = 0; i < items.length; i += 100) {
processChunk(items.slice(i, i + 100));
await new Promise(resolve => setImmediate(resolve));
}How to stop it coming back
- Set latency budgets for the largest supported input, not just average requests
- Monitor event-loop delay separately from downstream latency and worker-pool queues
- Load-test health checks while CPU-heavy endpoints run concurrently
- Review request-path synchronous APIs and unbounded serialisation
FAQ
Does making the function async fix CPU blocking?
No. Code before and between awaits still executes on the calling thread. Use a real asynchronous I/O API, partition the CPU work with scheduling boundaries, or move it to a bounded worker pool.
Will increasing UV_THREADPOOL_SIZE help?
Not for ordinary JavaScript loops or JSON.parse on the main thread. That setting affects operations using libuv’s worker pool. Confirm which resource queues the work before tuning it.
Why is my timeout callback also late?
The callback needs the event loop to run. A blocked loop cannot enforce its own timer promptly. Bound work proactively and consider an external supervisor for detecting a completely unresponsive process.
Related
Other errors engineers hit next to this one
- UnicodeDecodeError: 'utf-8' codec can't decode byte
- AssertionError: daemonic processes are not allowed to have children
- [CRITICAL] WORKER TIMEOUT (pid:1234)
- Mutable default argument retains state across calls
- CommitFailedException: Commit cannot be completed since the group has already rebalanced
- Consumer group stuck rebalancing — poll timeout has expired
- The same message processed twice (at-least-once delivery)
- Messages processed out of order across partitions