Linux / shell

Too many open files — a process limit or a descriptor leak?

Written and reviewed by Sahil Srivastav

File descriptorsSocketsResource lifecycle
OSError: [Errno 24] Too many open files

What this error actually means

On Linux, an open file descriptor can refer to a regular file, socket, pipe, event descriptor or other kernel object. The Python example reports EMFILE: this process cannot allocate another descriptor within its limit. A web service can therefore hit Too many open files without opening many ordinary files at all.

The failing open or accept operation is often the victim. A different request path may have leaked sockets for hours before a healthy request needs the final slot. The important distinction is whether descriptor use returns to a stable baseline after work completes. A bounded workload that legitimately reaches its configured ceiling needs capacity planning; a count that ratchets upwards needs ownership repaired.

A second failure, ENFILE, concerns the system-wide open-file table and is often worded Too many open files in system. Do not confuse a host-wide shortage with one service’s RLIMIT_NOFILE. Also, ulimit in your terminal reports the terminal’s limits, which need not match a process launched by systemd or a container runtime.

Causes, most common first

  1. 1A file or response body is not closed on every path. Happy-path cleanup can be skipped by exceptions, early returns or cancellation. Some HTTP clients retain a socket until the response is consumed or closed. Repeating a rare failure eventually consumes the process’s descriptors.
  2. 2Connections or watchers grow without a bound. Creating a new client or watcher for each request can keep sockets, notification descriptors and pools alive. Multiple individually bounded pools can still produce an unbounded process total when their owners are never removed.
  3. 3The configured limit is below measured legitimate demand. A service supporting many concurrent connections may have a stable descriptor count that still approaches its inherited soft limit. Account for listeners, logs, outbound sockets and internal descriptors as well as inbound clients before sizing headroom.
  4. 4Another workload exhausts host-wide capacity. If the actual errno is ENFILE and many services fail together, inspect the system-wide file table. Increasing only this process’s limit cannot create host capacity and may let it consume an even larger share.

When you see it

  • New connections fail while existing requests or sockets continue working
  • Restarting restores service but the failure returns after similar traffic volume
  • Descriptor count rises after cancelled requests and never returns to baseline
  • Changing ulimit in an SSH session has no effect on the running service

How to diagnose it

Step 1

Read the failing process’s actual limit

Replace 1234 with the live service PID. Run from a namespace where that PID is visible and with permission to inspect it. The limits file distinguishes the soft enforcement threshold from the hard ceiling.

app_pid=1234
cat /proc/"$app_pid"/limits

Step 2

Count and classify descriptors

Sample during load and again after requests drain. The symlink targets identify sockets, pipes, files and anonymous descriptors. Treat each count as a snapshot because the process can open or close descriptors while it is being inspected.

app_pid=1234
find /proc/"$app_pid"/fd -mindepth 1 -maxdepth 1 -type l | wc -l
ls -l /proc/"$app_pid"/fd

Step 3

Look at the objects held open

Where lsof is installed, filter to the affected process and examine repeated targets or socket families. Its rows are not an exact descriptor count because memory mappings and other entries can appear. /proc/PID/fd is the better count.

lsof -nP -p 1234

Step 4

Check the launch configuration and host table

For app.service, compare LimitNOFILE with /proc after a restart. The kernel counters help investigate a confirmed system-wide ENFILE event; do not treat them as a replacement for the individual process’s limits.

systemctl show app.service -p LimitNOFILE
cat /proc/sys/fs/file-nr
cat /proc/sys/fs/file-max

The fix

Give each descriptor a clear owner and scope its lifetime. Use context managers or finally blocks for files and client responses. Ensure cancellation and exception paths release resources too. Reuse a bounded network client where appropriate, and close its pool when the owning component shuts down.

Reproduce the suspected leak with repeated failed requests, then allow in-flight work to finish. The acceptance criterion is that descriptor count returns near the warm baseline, not that the next request happens to succeed. A larger limit can delay a leak for days while increasing its eventual impact.

If demand is bounded and measured, set an appropriate service or runtime nofile limit with operating headroom. Apply it at the real launch point and verify the new process’s /proc limits after rollout. A terminal ulimit command does not retroactively alter an unrelated running service.

When overload creates too many legitimate sockets, bound concurrency and idle connection retention as well. The file limit is a last boundary, not an admission-control strategy. Increasing it without enough memory, downstream capacity and cleanup discipline just moves the next failure elsewhere.

How to stop it coming back

  • Monitor open descriptors as a fraction of the process limit and watch the post-load baseline over time.
  • Include disconnects, timeouts and cancelled streams in lifecycle checks, since clean success paths rarely expose leaks.
  • Budget descriptor use across all pools in one process and all instances on a host.

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Why does ulimit -n show a different number?

It reports the current shell’s soft limit. Services inherit limits from their own launcher. Read /proc/PID/limits for the failing process and configure the service manager or runtime that actually starts it.

Can I close leaked descriptors from another process?

Do not try to repair ownership by deleting /proc entries or guessing descriptor numbers. Capture evidence, correct lifecycle code and recover the affected service through its normal restart procedure when needed.

Does a high socket count prove a leak?

No. Persistent client connections and bounded keep-alive pools can be legitimate. A leak is indicated by ownership evidence and growth that persists after the corresponding work should have released resources.

Related

Other errors engineers hit next to this one

Full error and symptom index →