Node.js
Node.js — ERR_UNHANDLED_REJECTION inside a worker thread
Written and reviewed by Sahil Srivastav
UnhandledPromiseRejection: This error originated either by throwing inside of an async function without a catch block, or by rejecting a promise which was not handled with .catch(). The promise rejected with the reason "worker task failed".What this error actually means
A worker has its own JavaScript execution environment and promise lifecycle. An async operation started inside it can reject without an owner just as it can on the main thread. With a fatal rejection policy, that failure becomes an uncaught worker exception and can terminate the worker. The parent needs to handle the Worker error event and settle every task affected by its loss.
The specific ERR_UNHANDLED_REJECTION wrapper commonly appears when the rejected value is not an Error, such as the string in the example. Rejecting an Error may display the original error instead. Search by task and worker identity rather than assuming every unhandled rejection has the same top-level code.
A parent awaiting a promise for a worker result is not automatically awaiting the worker’s internal async callback. postMessage sends data; it does not connect two promise chains. A reliable task protocol explicitly carries success or failure by task ID, while worker crash handling covers failures that prevent any result message from arriving.
Causes, most common first
- 1An async message listener has no local catch. parentPort.on("message", async (...) => ...) invokes an async function, but the event emitter does not inherently turn its returned rejection into a task result. A failure after an await escapes unless the listener catches it or explicitly supervises the returned promise.
- 2The worker starts detached helper work. The main task catches its own awaited operations but invokes a helper without awaiting it. The task can report success before that helper rejects. The worker then crashes after the parent believes the operation completed, making both diagnosis and replay decisions ambiguous.
- 3The parent handles messages but not worker lifecycle. A result listener alone cannot observe failure before the worker sends a message. The parent needs error and exit handling and a mapping of pending task IDs. Otherwise a crashed worker permanently consumes queue capacity through unresolved promises.
- 4Failure serialisation or reporting itself fails. Posting an unsupported value or building an error payload from unsafe application state can throw inside the catch path. Keep the task failure message small and serialisable. A failure while reporting failure belongs to worker crash handling, not an assumption that the parent received the result.
When you see it
- A worker exits after logging an async dependency failure
- The parent’s task promise waits forever because only successful messages resolve it
- An unhandled Worker error event also terminates the parent process
- Restarting the worker retries a task whose side effects may already have occurred
How to diagnose it
Step 1
Reproduce under an explicit rejection policy
Run a disposable worker task that rejects after an await. Record worker ID, task ID, error event and exit code. Workers normally inherit executable arguments unless configured otherwise, so inspect explicit worker execArgv overrides as well.
node --unhandled-rejections=strict --trace-uncaught main.mjsStep 2
Inspect the callback’s promise ownership
Follow the entire async message handler, including helpers started before posting success. Every required operation must be awaited before reporting completion. Catching a synchronous call to postMessage in the parent does not catch later work inside the worker.
Step 3
Audit parent settlement paths
List the events that can finish a task: result message, task error message, deadline, worker error and worker exit. Verify they converge on one idempotent settlement function that removes timers and map entries. Error and exit can both occur for the same worker failure.
Step 4
Check side-effect evidence before retrying
If the worker writes to a database or external API, look up the operation by a stable task key. A crash can happen after the write but before the result message. Blind replay after worker replacement can duplicate a successful operation.
The fix
Catch task-level failures inside the worker and send a structured error result with the task ID. Reject with Error objects for useful stack information, but transfer only the fields the parent needs. Preserve detailed internal diagnostics separately from public response messages.
In the parent, attach error and exit listeners before dispatching tasks. When a worker is lost, reject or requeue its pending tasks according to an explicit delivery contract. Settle each promise once and remove all associated state so a late message or second lifecycle event cannot complete it again.
Use a bounded worker pool and a bounded task queue. Do not create a worker per request to avoid lifecycle handling; that adds startup and memory costs while retaining the same failure ambiguity. Replace failed workers with restart backoff so a poison task does not create a rapid crash loop.
Make retries safe for side effects with durable task state, idempotency keys or outcome reconciliation. A worker error means execution did not return normally; it does not prove the operation did nothing. CPU-only deterministic tasks are much easier to replay than tasks combining computation and external writes.
// worker.mjs
import { parentPort } from "node:worker_threads";
parentPort.on("message", task => {
void executeAndReport(task).catch(reportingError => {
// Reporting itself failed: surface a worker failure to the parent.
setImmediate(() => { throw reportingError; });
});
});
async function executeAndReport({ id, input }) {
let result;
try {
result = await runTask(input);
} catch (error) {
parentPort.postMessage({
id, ok: false,
error: { message: error instanceof Error ? error.message : String(error) },
});
return;
}
parentPort.postMessage({ id, ok: true, result });
}
// Parent: handle both worker.on("error", ...) and worker.on("exit", ...).How to stop it coming back
- Test a rejection after await and a worker exit before any result message
- Assert pending task maps and timers return to baseline after worker failure
- Keep task result envelopes serialisable and separate task errors from worker crashes
- Document whether replay is safe after an uncertain side-effect outcome
FAQ
Does the main thread’s try/catch catch worker exceptions?
It catches synchronous errors in creating or messaging the worker, not later exceptions in the worker’s execution environment. Observe Worker lifecycle events and connect them to your own task promises.
Why do I receive both error and exit?
An uncaught worker exception can emit error and then terminate the worker, producing exit. Both describe the same failure lifecycle. Cleanup and task settlement must therefore be idempotent rather than rejecting and requeueing twice.
Should every task failure terminate its worker?
Expected input or dependency failures can be reported through the task protocol while a healthy worker continues. Unexpected failures that leave worker state uncertain should cause replacement. Do not silently continue after a crash-level invariant failure merely to avoid a restart.
Related
Other errors engineers hit next to this one
- Exactly-once claim fails at an external side effect
- Dead-letter queue growing without an alert
- Idempotency key reused with a different request body
- Read-after-write returned stale data from a replica
- Replication lag: replica served stale data after a write
- command not found in a script that works interactively
- Permission denied when executing a script
- bad interpreter: No such file or directory with CRLF