Concurrency
Thread-pool sizing: interview questions and how to answer them
A thread pool is sized from the work it performs and the resources it consumes, not from a memorised multiple of CPU cores.
Written and reviewed by Sahil Srivastav
What it actually is
A pool bounds concurrent work by owning a queue and a fixed or bounded set of workers. CPU-bound tasks compete for cores, while blocking tasks occupy workers while the CPU is idle; those two workloads therefore need different sizing models.
The useful size is the size at which throughput stops improving before queueing latency and downstream contention become unacceptable. Little’s Law gives a practical check: concurrency is throughput multiplied by time in system. A larger pool can hide slow work briefly, but it cannot create database connections or CPU cycles.
Why it matters in production
An oversized pool causes context switching, heap pressure, and database-pool contention. An undersized pool makes unrelated requests wait behind blocking work. The incident often appears as “the service has idle CPU but high latency” because all workers are sleeping on I/O.
Nested submission is especially dangerous: every worker can wait for work queued behind itself, producing starvation without a deadlock in the lock sense.
How it works
CPU-bound work
Start near the number of runnable cores and measure. Extra workers only add scheduling overhead once all cores are busy.
Blocking work
A pool serving blocking calls needs enough workers to cover the blocked fraction, but the downstream resource must be bounded too. A pool of 200 against a database pool of 20 creates 180 waiters.
Queue policy
An unbounded queue converts overload into latency and memory growth. A bounded queue plus rejection or backpressure makes saturation observable.
Separate workloads
Keep request handling, CPU work, and slow external calls in separate pools so one dependency cannot consume every worker.
Implementing it
Measure active workers, queue depth, task wait time, execution time, and rejection count. Change one limit at a time under a representative load test.
Set queue bounds from a latency budget. A request waiting longer than its deadline is no longer useful work.
Propagate cancellation and timeouts into blocking calls; interrupting a worker that remains blocked does not free capacity.
ExecutorService pool = new ThreadPoolExecutor(
8, 8, 0, TimeUnit.MILLISECONDS,
new ArrayBlockingQueue<>(100),
new ThreadPoolExecutor.CallerRunsPolicy());Interview questions and how to answer them
How would you size an I/O pool?
Begin with the measured service time and blocked fraction, then validate with a bounded load test. The pool must not exceed the downstream capacity, and queue wait must fit the request deadline.
Why can more threads reduce throughput?
They compete for cores, increase context switches and cache misses, and create more simultaneous downstream requests. Once the bottleneck is saturated, extra concurrency only lengthens the queue.
What does a bounded queue buy you?
It turns unlimited memory growth into a visible overload decision: reject, shed, or run the work in the caller. That lets backpressure reach the producer.
How do you spot pool starvation?
Workers are busy or waiting, queue depth rises, CPU may be low, and tasks that could run are queued behind tasks waiting for their own children or a depleted resource.
Answers that lose the round
- Using `2 * cores` for every workload
- Adding workers when the database pool is already saturated
- Leaving the queue unbounded
- Sharing one executor with scheduled, request, and blocking tasks
- Ignoring queue wait when calculating latency
- Submitting tasks from a worker and waiting synchronously for their children
FAQ
Should a pool have one thread per request?
No. Requests should be bounded by the work and dependency capacity they consume; one thread per request makes overload expensive and unpredictable.
Is a queue always bad?
No, but its bound and deadline must be deliberate. A short queue absorbs bursts; an infinite queue hides a permanent overload.
How many pools should a service have?
Enough to isolate materially different blocking and scheduling behaviour, while keeping the number small enough to observe and tune.