Node.js
Node.js — Reached heap limit, Allocation failed
Written and reviewed by Sahil Srivastav
FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memoryWhat this error actually means
V8 could not allocate another JavaScript object within its heap budget after attempting garbage collection. The allocation named near the fatal output is the point where available space ran out; it does not identify the objects that occupied the heap. An innocent JSON response can fail because a long-lived registry retained thousands of previous requests.
Separate JavaScript heap from total process memory. Buffer backing stores and other native allocations contribute to resident memory without all appearing in heapUsed. A process killed by its container may leave no V8 fatal message at all. Conversely, V8 can hit its own configured ceiling while the container still has memory available. These are different limits with different evidence.
The useful distinction is retained growth versus temporary amplification. Retained growth survives quiet periods and collections. A large export can create a sudden peak because query rows, transformed objects and the serialised string coexist. Both exhaust the same heap, but only the first requires finding an owner that never releases references.
Causes, most common first
- 1A long-lived collection has no eviction rule. Maps keyed by request IDs, memoised results and per-user registries grow with the number of distinct inputs. Expiring a database row does not remove its corresponding in-memory reference. Inspect the owner and its deletion path before blaming the class occupying the most bytes.
- 2The request materialises more than one full representation. Fetching every row, mapping into new objects and then JSON.stringify create overlapping lifetimes. Promise.all across many such requests multiplies the peak. A maximum response size enforced after serialisation arrives too late to protect the heap.
- 3Listeners or pending work retain completed requests. A closure attached to a global emitter can retain the request, authentication context and payload. A queue waiting behind a slow downstream service has the same effect even when every item will eventually complete: admission exceeds the rate at which memory is released.
- 4The legitimate working set exceeds the configured budget. This is plausible after a deliberate increase in bounded cache size or dataset size. Establish that object counts stabilise under sustained load first. Raising a limit without measuring the live set can turn a diagnosable V8 failure into an abrupt container kill.
When you see it
- Heap usage repeatedly recovers to a higher baseline after traffic settles
- Garbage collection consumes more CPU while completed requests per second fall
- One large tenant or export crashes an otherwise stable service
- A restart restores capacity but the same workload eventually reproduces the failure
How to diagnose it
Step 1
Sample the relevant memory categories inside the service
Record these fields together with in-flight requests and queue depth. RSS includes more than V8; external and arrayBuffers help identify buffer-heavy traffic. Running the command in a separate process only demonstrates the API, so insert the same sampling into the affected process.
node -e "console.log(process.memoryUsage()); console.log(require('node:v8').getHeapStatistics().heap_size_limit)"Step 2
Capture a snapshot before the final allocation failure
For a disposable diagnostic instance with spare memory and protected storage, enable a near-limit snapshot. Snapshot generation can pause the process and require substantial extra memory; avoid triggering it simultaneously on every production replica. Snapshots may contain secrets and request data.
node --heapsnapshot-near-heap-limit=1 server.jsStep 3
Compare retained paths under the same workload
Compare snapshots before and after repeated requests, then examine retaining paths for growing groups. A large array is a clue; the module cache, emitter or map keeping it reachable is the actual repair point. Repeat after allowing queues to drain so legitimate outstanding work is not mistaken for a permanent leak.
Step 4
Measure the largest individual operation
Replay a production-shaped export in staging at concurrency one, then at the intended concurrency limit. If peak memory scales with result size but returns to baseline, bound the operation rather than searching indefinitely for a leak.
The fix
Give every process-wide cache a size or weight budget, an expiry policy where appropriate and metrics for entries and evictions. A TTL alone does not constrain a burst of new keys within the TTL. Remove listeners and request registrations in a cleanup path that runs after success, failure and cancellation.
Push pagination into the database query and stream output with backpressure. Limit the number of simultaneous exports. Replacing an array with a stream has no benefit if the next function collects every chunk before forwarding it.
For queued tasks, reject or defer admission once the queue reaches a measured bound. Store identifiers rather than complete request objects, and propagate cancellation so abandoned requests stop retaining downstream work.
Increase --max-old-space-size only after the retained set is demonstrably bounded. Leave room below the process or container memory ceiling for native allocations, thread stacks and workload peaks; V8’s heap limit is not a reservation for all process memory.
How to stop it coming back
- Exercise the largest supported payload and tenant dataset, not only average fixtures
- Graph settled heap baseline alongside RSS, concurrent requests and queue length
- Check cancellation and exception paths for retained listeners and work items
- Keep a snapshot runbook with storage, access controls and a safe diagnostic replica
FAQ
Does doubling the heap solve this?
It can accommodate a measured, bounded working set. It cannot make an unbounded map or request safe. Compare the settled live set before and after sustained traffic; growth proportional to lifetime traffic predicts another failure at the larger ceiling.
Why is RSS high when heapUsed looks reasonable?
Resident memory also contains native allocations, buffer backing stores, code and allocator overhead. Start with external and arrayBuffers, then investigate native memory if the categories do not explain the gap. A V8 heap snapshot alone does not account for the entire process.
Can a catch block recover from this fatal error?
Do not design recovery around catching V8’s fatal heap exhaustion. Protect the process by bounding work before allocation, and let a supervisor replace failed instances. Retrying the same oversized request immediately on every replica only distributes the crash.
Related
Other errors engineers hit next to this one
- ThreadLocal value leaking across requests on a pooled thread
- Consumer stuck in Object.wait() with work already in the queue
- java.lang.IllegalMonitorStateException: current thread is not owner
- Thread pool starvation — every worker waiting on a task in its own pool
- Partially constructed object published by double-checked locking
- Lost update from get-then-put on a ConcurrentHashMap
- CompletableFuture failed with nothing logged
- awaitTermination never returns and the JVM will not exit