Concurrency
A CompletableFuture exception swallowed silently
Written and reviewed by Sahil Srivastav
# The stage failed. Nothing was logged. This is what you get when you finally ask:
java.util.concurrent.CompletionException: java.lang.IllegalStateException: order 88213 has no shipment
at java.base/java.util.concurrent.CompletableFuture.encodeThrowable(CompletableFuture.java:315)
at java.base/java.util.concurrent.CompletableFuture$UniApply.tryFire(CompletableFuture.java:646)
at java.base/java.util.concurrent.CompletableFuture$Completion.run(CompletableFuture.java:482)
Caused by: java.lang.IllegalStateException: order 88213 has no shipment
at com.example.ship.ShipmentStage.apply(ShipmentStage.java:41)
... 6 more
# Note: no frame from the code that started the pipeline. The submitter is not in the trace.What this error actually means
A `CompletableFuture` that fails does not throw anywhere. The exception is *captured* as the completion value of that future and then propagated down the dependent chain as a result, not as a control-flow event. If nobody ever asks the last future in the chain what happened — no `join`, no `get`, no `exceptionally`, no `whenComplete` — the exception is stored in an object that becomes garbage, and the failure leaves no trace at all.
This is a deliberate design choice and it is the opposite of how a thread-based task behaves. An exception escaping a `Runnable` on a raw thread reaches the thread’s uncaught-exception handler and gets printed. An exception inside a `CompletableFuture` stage never escapes: the framework catches it precisely so it can be delivered to whoever asks. There is no default handler in the path because, from the framework’s point of view, nothing went uncaught.
The gap widens with `thenApply` chains whose returned future is discarded. Calling `future.thenApply(this::render)` and ignoring the result creates a new future that will hold the failure of either stage and be immediately unreferenced. The work still runs, the error still occurs, and the return value you dropped was the only handle to the outcome.
Two further details make the trace harder to read when you do finally see it. Exceptions are wrapped — `CompletionException` from `join`, `ExecutionException` from `get` — so the type you catch is not the type that was thrown, and naive `catch (IllegalStateException e)` around a `join` does not match. And the stack trace contains the async stage frames but not the frames of the code that *started* the pipeline, because that code is no longer on any stack.
Causes, most common first
- 1The returned future is discarded. The most common shape by far. `thenApply`, `thenCompose`, and `thenAccept` all return a *new* future carrying the outcome; ignoring the return value throws away the only reference through which a failure could ever be observed. Static analysis flags this as an ignored return value, which is why enabling that check pays for itself immediately.
- 2No terminal error handler on the chain. The pipeline ends in `thenAccept` that performs a side effect, with no `exceptionally` or `whenComplete` after it. A failure anywhere upstream short-circuits every remaining `then*` stage — they are skipped, not executed with a null — and stops at the end of the chain with nobody listening.
- 3`exceptionally` placed in the middle of the chain. It only observes failures from stages *above* it. Anything that fails below is unhandled again. The handler must be last, or you need one at the end in addition to any recovery stages in the middle.
- 4Fire-and-forget `runAsync` for background work. A stage submitted purely for its side effect, with the future never retained. This is the pattern that loses outbox writes, cache invalidations, and audit records — the ones you discover are missing days later during reconciliation.
- 5Exception thrown inside a completion callback. A `whenComplete` or `thenAccept` body that itself throws produces a failure on the *downstream* future. If the handler is the last stage, its own exception is swallowed by the same mechanism it was added to prevent.
- 6Failure inside `allOf` attributed to nothing. `allOf(...)` completes exceptionally with one of the failures if any input fails, and the others are never inspected. Logging only the `allOf` result loses every failure but one, which makes multi-leg fan-out failures look like single flukes.
- 7Handling the wrong exception type. `join` wraps in `CompletionException`, `get` in `ExecutionException`. A `catch` on the domain exception does not match, so the handler is skipped and the exception propagates somewhere generic — or is caught by an over-broad handler and logged at debug level.
When you see it
- A request returns 200 with a partial or default response and no error anywhere in the logs
- Work that should have happened simply did not: no email sent, no outbox row written, no index updated
- Success and failure metrics both fail to increment for the affected operations, so the gap is only visible as missing volume
- Adding a single `.join()` while debugging makes an exception appear immediately that was never logged before
- A `catch` block around a `join` does not match the exception type you expected, because it arrives wrapped
- Stack traces from async failures lack any frame from the calling code, making them hard to attribute to an endpoint
- A `timeout` or `orTimeout` fires and the resulting `TimeoutException` is likewise never observed
How to diagnose it
Step 1
Add a terminal handler and let the failures surface
Before theorising, attach a `whenComplete` that logs at the end of every suspect chain. Failures that were invisible for weeks typically appear in the first minute. This is a diagnostic that is also most of the fix.
future.whenComplete((v, ex) -> {
if (ex != null) log.error("stage failed", ex);
});Step 2
Turn ignored return values into build failures
This bug has a mechanical signature: a `CompletionStage` method whose result is discarded. Error Prone’s `FutureReturnValueIgnored` check finds every instance across the codebase, which is far more complete than reading code.
./gradlew compileJava -PerrorProneChecks=FutureReturnValueIgnoredStep 3
Check whether the work ran at all
Distinguish "stage failed silently" from "stage never ran". Log entry and exit of each stage body with the same correlation id; an entry with no exit is a failure inside, and no entry at all means the chain short-circuited further upstream.
Step 4
Unwrap before you inspect
When you do catch something, unwrap the completion wrapper or you will log the wrapper type and lose the ability to branch on the real cause.
Throwable root = (ex instanceof CompletionException || ex instanceof ExecutionException)
? ex.getCause() : ex;Step 5
Establish which executor the stages ran on
Non-`Async` continuations run on whichever thread completed the previous stage, and `*Async` without an executor uses the common ForkJoinPool. Log the thread name inside each stage: work landing on `ForkJoinPool.commonPool-worker-*` means you are sharing a JVM-wide pool with library code and with parallel streams.
jcmd <pid> Thread.print | grep -c "ForkJoinPool.commonPool-worker"The fix
Terminate every chain with a handler, and make that a rule rather than a habit. A final `whenComplete` that logs the throwable and increments a failure counter costs one line and converts an invisible failure class into an observable one. `whenComplete` is the right choice for the terminal position because it observes both outcomes and, unlike `handle`, does not require you to invent a substitute value.
Never discard the future returned by a `then*` method. Either return it to the caller so the failure becomes their problem to handle, or attach the terminal handler right there. An expression-statement call to `thenApply` is always a bug.
Choose between the three handlers on purpose. `exceptionally` recovers from a failure by supplying a fallback value and leaves success untouched — use it where a default is genuinely correct. `handle` sees both outcomes and must produce a value, so use it when the substitution depends on which happened. `whenComplete` observes both and changes neither, which is what you want for logging, metrics, and cleanup.
Always pass an explicit executor to `supplyAsync`, `thenApplyAsync`, and the rest. Defaulting to the common ForkJoinPool means your blocking work competes with parallel streams and library internals on a pool whose parallelism is one less than the processor count, and it is a direct route to the starvation failure where every worker waits on a stage that needs a worker.
Unwrap `CompletionException` and `ExecutionException` at the boundary before mapping errors to responses, and never catch the wrapper and log only its message. Carry the cause into whatever domain error your API returns, so the client-facing message and the logged trace describe the same failure.
For fan-out, inspect each leg individually rather than relying on `allOf`. Attach a handler per future so every failure is recorded, then use `allOf` purely as the barrier. Otherwise a fan-out where three of five legs fail is indistinguishable from one where one did.
For genuinely fire-and-forget work, decide explicitly whether losing it is acceptable. If it is not — an outbox write, an audit record — it does not belong in an unretained future at all; it belongs in a durable queue or a transactional outbox that survives process death as well as exceptions.
// Silent: the returned future holds the failure and is immediately garbage
CompletableFuture.supplyAsync(() -> loadOrder(id))
.thenApply(this::attachShipment); // may throw; nobody asks
// Silent: exceptionally is not last, so render() failures are unobserved
future.exceptionally(ex -> Order.EMPTY)
.thenApply(this::render);
// Observed: explicit executor, terminal handler, unwrapped cause
CompletableFuture<Shipment> f =
CompletableFuture.supplyAsync(() -> loadOrder(id), ioPool)
.thenApplyAsync(this::attachShipment, cpuPool)
.whenComplete((v, ex) -> {
if (ex != null) {
log.error("shipment pipeline failed for order={}", id, unwrap(ex));
failures.increment();
}
});
static Throwable unwrap(Throwable t) {
return (t instanceof CompletionException || t instanceof ExecutionException)
? t.getCause() : t;
}
// Fan-out: handle each leg, then use allOf only as the barrier
List<CompletableFuture<Leg>> legs = ids.stream()
.map(i -> fetch(i, ioPool).whenComplete((v, ex) -> recordLeg(i, ex)))
.toList();
CompletableFuture.allOf(legs.toArray(CompletableFuture[]::new))
.whenComplete((v, ex) -> finish(legs));How to stop it coming back
- Enable Error Prone’s `FutureReturnValueIgnored` (or an equivalent SpotBugs rule) and treat it as an error, not a warning
- Require a terminal `whenComplete` on every chain in review; the absence of one is the bug, not a style choice
- Ban defaulted executors in async stages so pool ownership is always explicit and greppable
- Emit a metric from the terminal handler so async failures show up on dashboards rather than only in logs
- Unwrap completion exceptions in one shared helper, so no boundary logs the wrapper and loses the cause
- Move any async side effect you cannot afford to lose out of fire-and-forget futures and into a durable outbox
FAQ
Why is there no uncaught-exception handler for this?
Because nothing goes uncaught. The framework catches the exception on purpose and stores it as the future’s completion value so it can be delivered to a dependent stage or to a caller who joins. From the runtime’s perspective the task completed normally — with a failure as its result — so the thread’s uncaught handler is never involved.
What is the difference between `exceptionally`, `handle` and `whenComplete`?
`exceptionally` runs only on failure and returns a replacement value. `handle` runs on both outcomes and must return a value, so it can transform either. `whenComplete` runs on both, returns nothing, and passes the original outcome through unchanged — which makes it the correct terminal observer for logging and metrics, since it cannot accidentally convert a failure into a success.
Do the downstream stages run with a null value when a stage fails?
No. A failure short-circuits: every dependent `then*` stage is skipped entirely and the failure propagates to the end of the chain. Only handlers that explicitly accept a throwable are invoked. This is why a missing side effect, rather than a `NullPointerException`, is the usual symptom.
Which exception type will I actually catch?
`join()` wraps the cause in an unchecked `CompletionException`; `get()` wraps it in a checked `ExecutionException`. Cancellation surfaces as `CancellationException` from both. Always unwrap with `getCause()` before matching on a domain type — catching the domain exception directly around a `join` silently never matches.
Does the failure still get lost if I use virtual threads?
Yes, if you keep using `CompletableFuture` the same way — the capture-and-store mechanism is unchanged. What virtual threads enable is structured concurrency, where a scope owns its subtasks, propagates their failures to the joining code, and cannot let a child fail unobserved. That removes the failure class rather than making it easier to log.
Related
Other errors engineers hit next to this one
- Process exits before asynchronous writes finish
- ERR_MODULE_NOT_FOUND during ESM/CommonJS migration
- Unbounded JSON body blocks the loop or exhausts memory
- Undici / fetch connections remain occupied
- ERR_UNHANDLED_REJECTION in a worker thread
- command not found in a script that works interactively
- Permission denied when executing a script
- bad interpreter: No such file or directory with CRLF