Python
Task was destroyed but it is pending!
Written and reviewed by Sahil Srivastav
Task was destroyed but it is pending!
task: <Task pending name='Task-14' coro=<flush_metrics() running at /app/metrics.py:61> wait_for=<Future pending cb=[Task.task_wakeup()]>>What this error actually means
This is asyncio telling you that work disappeared. A `Task` object was finalised while its coroutine had not finished, so whatever remained after the last `await` will never execute — no exception is raised anywhere, no caller is notified, and the log line is the only trace. It is a warning by severity and a silent-data-loss report by consequence.
There are two ways to get here and they are worth separating. The first is garbage collection: `asyncio.create_task` returns a task the loop only holds a weak reference to while it is waiting, so if your code keeps no strong reference, the task can be collected mid-flight at an arbitrary point. This is the one that produces genuinely random, low-frequency loss in production and is close to impossible to reproduce on demand.
The second is shutdown. `asyncio.run` cancels remaining tasks when the main coroutine returns, and a task cancelled while suspended never resumes. If that task was flushing metrics, committing an outbox row, acking a message or closing a file, the effect is the same as the first case: partially applied work with no error.
Notice what the message gives you for free — the coroutine name and the exact source line it was suspended at. That line tells you precisely which half of the operation completed and which half did not.
Causes, most common first
- 1create_task with the result discarded. The canonical cause. `asyncio.create_task(send_webhook(...))` with no variable assignment and no awaiting owner. The loop does not keep the task alive on your behalf while it waits, so the garbage collector is free to take it. The code reads as "run this in the background" and behaves as "run this, probably".
- 2Shutdown that does not drain in-flight work. The main coroutine returns, or SIGTERM arrives, and pending tasks are cancelled where they stand. Very common in containers, where a rolling deploy cancels mid-operation tasks on every pod several times a day.
- 3A task owned by an object that goes out of scope. A connection handler, websocket session or per-request service stores its task on `self`, and then the object itself becomes unreachable when the request ends. The strong reference exists but is reachable from nothing.
- 4Tasks awaiting something that will never complete. A future nobody resolves, a queue that receives no more items, a lock whose holder died. These stay pending until shutdown destroys them, so the warning is the first and only sign the task was stuck for hours.
- 5Loop replaced or closed while tasks remain. Multiple `asyncio.run` calls, or `new_event_loop` in the middle of a program, abandon whatever the previous loop still had scheduled. This pairs with `RuntimeError: Event loop is closed` in the same log window.
When you see it
- A small, irreproducible fraction of background writes never land, with no error anywhere
- The warning appears in bursts at shutdown, deploy, or SIGTERM-triggered restarts
- Frequency changes when unrelated code changes allocation patterns, because collection timing changed
- A handler returns success to the client while its follow-up work did not happen
- The named coroutine is always something fire-and-forget: audit logging, metrics, cache warm, webhook dispatch
How to diagnose it
Step 1
Read the suspend point in the message
The `running at <file>:<line>` fragment is the exact statement the coroutine was parked on. Everything before it ran; everything after it did not. That alone usually identifies the lost side effect.
Step 2
Enumerate pending tasks before shutdown completes
Log what is still in flight at the start of shutdown. If the list is non-empty every time, you have no drain phase and the warning is guaranteed, not incidental.
pending = [t for t in asyncio.all_tasks() if t is not asyncio.current_task()]
logger.warning("draining %d tasks: %s", len(pending), [t.get_name() for t in pending])Step 3
Run with development mode to get creation tracebacks
Debug mode records where each task was created and where exceptions were never retrieved, which turns "some coroutine vanished" into a file and line.
PYTHONASYNCIODEBUG=1 python -X dev -W error::RuntimeWarning app.pyStep 4
Audit every unassigned create_task
A mechanical search finds the fire-and-forget sites. Any `create_task` whose result is not assigned, gathered, or added to a set is a candidate for collection mid-flight.
grep -rn 'create_task(\|ensure_future(' --include='*.py' . | grep -v '=' The fix
Give every task an owner. The modern answer is `asyncio.TaskGroup` (Python 3.11+): tasks created inside the `async with` block are awaited at exit, exceptions propagate to the owner instead of vanishing, and a failure cancels siblings. It makes structured concurrency the default and removes both the collection hazard and the silent-failure hazard in one construct.
If you genuinely need work to outlive the request, keep a strong reference in a module-level `set` and discard it in a done callback: `task = asyncio.create_task(coro); tasks.add(task); task.add_done_callback(tasks.discard)`. This is the documented pattern for a reason — without it the task is collectable while suspended.
Build a real drain phase into shutdown. On SIGTERM, stop accepting new work, then await the outstanding tasks with a deadline, then cancel whatever is left and await the cancellations. In FastAPI or Starlette this belongs in the lifespan shutdown; in Kubernetes it needs `terminationGracePeriodSeconds` longer than the drain deadline, or the kernel kills you mid-drain regardless of what your code does.
Make the work durable rather than in-memory where losing it matters. If the follow-up is a webhook, a ledger write or an email, the reliable structure is to commit an intent row in the same transaction as the business change and have a separate consumer act on it. Then a cancelled task costs a delay instead of a lost event, because the record of what should happen survives the process.
Never swallow the completion of a task you started. A task whose exception is never retrieved logs `Task exception was never retrieved` at garbage-collection time — the same class of invisible failure as this warning, and the reason background errors can go unnoticed for months.
# Collectable while suspended; also loses any exception it raises
async def handle(request):
asyncio.create_task(flush_metrics(request))
return web.Response(text="ok")
# Owned, awaited, and exceptions surface at the owner
async def handle(request):
async with asyncio.TaskGroup() as tg:
tg.create_task(flush_metrics(request))
tg.create_task(write_audit(request))
return web.Response(text="ok")
# When the work must outlive the caller, hold a strong reference
_background: set[asyncio.Task] = set()
def spawn(coro) -> asyncio.Task:
task = asyncio.create_task(coro)
_background.add(task)
task.add_done_callback(_background.discard)
return task
# Drain with a deadline on shutdown
async def shutdown() -> None:
if _background:
done, pending = await asyncio.wait(_background, timeout=20)
for task in pending:
task.cancel()
await asyncio.gather(*pending, return_exceptions=True)How to stop it coming back
- Prefer `TaskGroup` over bare `create_task` everywhere; make unassigned `create_task` a lint failure
- Give tasks names (`create_task(coro, name="webhook-dispatch")`) so this warning identifies the subsystem instead of printing `Task-14`
- Make the shutdown drain a tested code path: send SIGTERM under load in CI and assert no pending-task warnings and no lost writes
- Set the container grace period above your drain deadline; a correct drain that gets SIGKILLed is still data loss
- For anything a user would notice losing, persist the intent transactionally rather than relying on an in-process task
Practise this failure in a real repository
Gronex ships the durability version of this failure: a service that commits a business change and dispatches the follow-up effect in memory, so a restart loses events. The test suite kills the process mid-dispatch and asserts every committed change eventually produces exactly one effect — which only a transactional outbox satisfies.
FAQ
Is this warning harmful or just noise?
It reports work that did not finish. Whether that matters depends entirely on what the coroutine was doing — nothing, for a cache warm; a lost payment notification, for a webhook dispatch. Treat it as a defect until you have read the coroutine named in the message.
Why is it so rare and unreproducible?
For the garbage-collection variant, the trigger is a collection cycle happening while the task is suspended, which depends on allocation pressure across the whole process. That is why an unrelated change to request volume or payload size moves the failure rate, and why a load test rarely shows it.
Does awaiting the task defeat the purpose of running it in the background?
Not if you await at the right boundary. `TaskGroup` runs tasks concurrently and only joins at the end of the block, so you keep the parallelism and lose only the ability to let work escape the block — which is precisely the part that was unsafe.
How does this relate to "Task exception was never retrieved"?
Same root cause, different outcome: an unowned task that failed rather than one that never finished. Both are consequences of creating tasks nothing is responsible for, and both disappear under structured concurrency.
Related
Other errors engineers hit next to this one
- FATAL ERROR: Reached heap limit Allocation failed
- Unhandled promise rejection crashes the process
- ECONNRESET: socket hang up on a reused connection
- MaxListenersExceededWarning: possible EventEmitter memory leak
- EADDRINUSE: address already in use
- ERR_HTTP_HEADERS_SENT
- Event loop blocked by synchronous work
- pg client already connected or released twice