Python
Celery WorkerLostError: SIGKILL or OOM
Written and reviewed by Sahil Srivastav
celery.exceptions.WorkerLostError: Worker exited prematurely: signal 9 (SIGKILL) Job: 7.What this error actually means
Celery’s parent process knows a child disappeared but cannot see the original Python exception when the operating system kills it. `WorkerLostError` is therefore a report about process ownership, not a root cause. Signal 9 means no cleanup or traceback was possible.
The common path is cgroup OOM: one task materialises a large result, the kernel selects the child, and Celery reports the loss. A hard time limit sends a signal too, while a native extension crash may leave signal 11. The remedy depends on the signal and the kernel event.
Retries make this worse when the task is deterministic and memory hungry: every retry starts another child, while the original partial side effect may already exist.
Causes, most common first
- 1Container or host OOM. The child exceeded its cgroup limit or the host had no reclaimable memory.
- 2Celery hard time limit. The worker is terminated after the configured deadline and the parent observes a lost child.
- 3Native extension crash. A C extension can terminate the process without a Python traceback.
- 4Forced deployment or node termination. Supervisor and orchestration signals can look like application failures.
When you see it
- WorkerLostError names SIGKILL or exit code 137
- The parent remains healthy and replaces children
- Tasks retry until the queue grows
- Kernel or container events show OOM kills
- Only one task type causes losses
How to diagnose it
Step 1
Check signal and exit code
Signal 9/137 strongly suggests an external kill; do not debug a missing Python traceback first.
kubectl describe pod <pod> | sed -n '/Last State:/,/Events:/p'
docker inspect <container> --format '{{.State.ExitCode}} {{.State.OOMKilled}}'Step 2
Read kernel evidence
The kernel records the decisive OOM event on the node or container runtime.
dmesg -T | rg -i 'out of memory|killed process|oom'Step 3
Correlate task memory
Log task name, input size, pid, and RSS at start and end so one payload can be compared with the kill time.
ps -o pid,rss,etime,cmd -C celeryThe fix
Fix the memory shape: stream files, chunk database reads, and bound per-task input.
Set Celery `worker_max_tasks_per_child` to recycle fragmentation-prone workers, but do not treat recycling as a leak fix.
Align soft and hard time limits with the real SLA and ensure the task is idempotent before enabling retries.
Reserve memory for the parent and sidecars when setting cgroup limits; the child does not own the whole limit.
Capture partial progress and make external writes deduplicate by task or business key.
@app.task(acks_late=True, autoretry_for=(TransientError,), retry_backoff=True)
def import_partitions(keys):
for key in keys: # bounded work per child
process_one(key)
# celeryconfig.py
worker_max_tasks_per_child = 100
task_soft_time_limit = 300
task_time_limit = 330How to stop it coming back
- Alert on worker replacement rate and exit code 137
- Load-test the largest real payloads under the production cgroup limit
- Keep retries bounded with a dead-letter policy
- Use task-level memory budgets
- Record whether a task crossed an external side-effect boundary before retrying
FAQ
Should I increase the memory limit?
Only after measuring a bounded workload and reserving headroom. If input size is unbounded, a larger limit delays the same failure.
Why is there no traceback?
SIGKILL cannot be caught by Python, so the process has no opportunity to print one.
Does max_tasks_per_child fix OOM?
It limits accumulated fragmentation or leaks between tasks; it cannot make one oversized task fit.
Related
Other errors engineers hit next to this one
- ECONNRESET: socket hang up on a reused connection
- MaxListenersExceededWarning: possible EventEmitter memory leak
- EADDRINUSE: address already in use
- ERR_HTTP_HEADERS_SENT
- Event loop blocked by synchronous work
- pg client already connected or released twice
- Sequelize / Knex pool acquire timeout
- ERR_STREAM_PREMATURE_CLOSE during an upload