Java / JVM

java.lang.OutOfMemoryError: unable to create new native thread

Written and reviewed by Sahil Srivastav

JVMThreadsOS limits
java.lang.OutOfMemoryError: unable to create new native thread
	at java.base/java.lang.Thread.start0(Native Method)
	at java.base/java.lang.Thread.start(Thread.java:809)

What this error actually means

The heap is irrelevant here. This message means the JVM asked the operating system to create a thread and the OS said no. Every Java thread is a native OS thread with its own stack reserved outside the heap — typically 512 KB to 1 MB — so threads consume native memory and a process-level or system-level thread slot.

The refusal comes from one of three ceilings: the process hit a limit on threads or process ids, the system ran out of native memory to reserve another stack, or a container cgroup `pids.max` was reached. Confusingly, raising `-Xmx` makes this *more* likely, since a larger heap leaves less address space and memory for thread stacks.

In a service that creates threads through bounded pools, this error is essentially impossible. So seeing it is strong evidence that something creates threads without bound — per request, per connection, per retry, or per reload — and never reaps them.

Causes, most common first

  1. 1A thread created per request or per task. Calling `new Thread(...).start()` inside a handler, or `Executors.newSingleThreadExecutor()` per operation without shutting it down. Each one leaks a live thread. This is the most common cause by a wide margin.
  2. 2An unbounded cached thread pool under load. `Executors.newCachedThreadPool()` creates a thread whenever no idle one is free and has no maximum. Combined with tasks that block on I/O, it grows to match arrival rate. It behaves perfectly until the day the downstream service slows down.
  3. 3Executors never shut down. A pool created in a component that is discarded but never has `shutdown()` called keeps its core threads alive forever. Common in reloadable modules, tests, and per-tenant service objects.
  4. 4A low pids or nproc ceiling. Container `pids.max`, Kubernetes pid limits, or `ulimit -u` set conservatively. Here the thread count may be entirely reasonable and the limit is the bug — but verify the count before believing that.
  5. 5Native memory exhausted by stack reservations. Thousands of threads at 1 MB of reserved stack each is gigabytes of native memory. With a large heap in a memory-capped container, the reservation fails even though no explicit thread limit was hit.

When you see it

  • Thread count grows monotonically with traffic and never falls during idle periods
  • The process RSS climbs while heap usage stays flat
  • It fires in whichever code happens to start a thread next, so the trace names an innocent component
  • New connections and scheduled tasks fail while existing request handling continues briefly
  • It appears in a container but not on a developer machine, because the cgroup pids limit is far lower

How to diagnose it

Step 1

Count threads and group them by name

Thread names reveal the leaking factory immediately, because a leak produces hundreds of threads with the same prefix and an incrementing counter.

jcmd <pid> Thread.print | grep -oP '^"[^"]+' | sed 's/[0-9]*$//' | sort | uniq -c | sort -rn | head

Step 2

Compare against the actual ceilings

Establish whether you are near a limit or far below it. In a container the cgroup value matters, not the host value.

cat /proc/<pid>/status | grep Threads
ulimit -u
cat /sys/fs/cgroup/pids.max

Step 3

Check whether threads are blocked or just numerous

A thread dump showing hundreds of threads parked on the same socket read or the same lock tells you the pool is growing because work is not completing — the fix is downstream, not in the pool.

jcmd <pid> Thread.print > threads.txt; grep -c "java.lang.Thread.State" threads.txt

Step 4

Account for native memory

Native Memory Tracking attributes reserved and committed bytes to thread stacks, which confirms whether stack reservation is what actually failed.

jcmd <pid> VM.native_memory summary

The fix

Replace unbounded creation with a bounded, named pool. A fixed pool with a bounded queue and an explicit rejection policy converts an unbounded resource leak into visible, handleable backpressure — which is what you want, because rejecting work loudly beats dying silently.

Own every executor’s lifecycle. If a component creates a pool, it must shut it down, and the shutdown must run on the failure path too. In a container framework, let the framework manage the pool so its lifecycle is tied to the component.

Name your thread factories. `order-dispatch-%d` instead of `pool-7-thread-%d` turns a future incident from an hour of archaeology into one line of a thread dump. This is the cheapest observability change available.

If threads are numerous because they are all blocked on I/O, the real fix is a timeout. An unbounded pool of threads waiting forever on a hung dependency is a resource leak caused by a missing timeout. Set connect and read timeouts on every outbound call.

Only if the count is genuinely justified should you raise limits — `pids.max` and `ulimit -u` — and reduce `-Xss` to shrink per-thread stack reservation. For workloads that are thread-per-request and I/O-bound, virtual threads remove the ceiling properly by decoupling tasks from OS threads.

How to stop it coming back

  • Ban `new Thread()` and `newCachedThreadPool()` in application code via a lint rule; require a bounded, named pool
  • Alarm on thread count trend — it is monotonic in a leak and therefore trivially detectable
  • Set timeouts on every outbound call. Most thread leaks are really timeout leaks
  • Include thread count in the health endpoint so a leak is visible before it becomes an outage
  • Load-test with a slow downstream dependency, not a fast one; that is the condition under which unbounded pools explode

Practise this failure in a real repository

Gronex ships executor-lifecycle failures as runnable repositories, including the starvation deadlock where pooled tasks submit sub-tasks to their own pool. The tests assert the pool makes progress under load rather than checking a happy path.

FAQ

Why is this an OutOfMemoryError if the heap is fine?

`OutOfMemoryError` covers any resource the JVM could not obtain, not just heap. Here the missing resource is a native thread — an OS-level allocation. The name is genuinely misleading, and it sends most people to heap analysis for no reason.

Will increasing -Xmx help?

It typically makes things worse. Thread stacks live in native memory outside the heap, so a bigger heap leaves less room for stacks in a memory-capped process or container.

How many threads is too many?

For a CPU-bound service, more than a small multiple of core count is already wrong. For thread-per-request I/O-bound services, hundreds can be legitimate and thousands almost never are. The number that matters is whether it is stable or growing.

Do virtual threads eliminate this?

For the blocking-I/O case, largely yes — millions of virtual threads map onto a small carrier pool, so the OS thread ceiling stops being the constraint. They do not help if the real bug is an executor you never shut down, and they do not remove the need for timeouts.

Related

Other errors engineers hit next to this one

Full error and symptom index →