Performance
Backpressure: interview questions and practical design
Backpressure keeps producers from creating work faster than consumers can safely process it by bounding queues, slowing admission, or rejecting work.
Written and reviewed by Sahil Srivastav
What it actually is
Backpressure keeps producers from creating work faster than consumers can safely process it by bounding queues, slowing admission, or rejecting work.
Without a bound, latency and memory grow together until the process or dependency fails. A queue only moves the overload unless its capacity and policy are explicit.
The useful interview answer is precise about the boundary: A bounded queue has a full condition that must map to a policy: block, drop, reject, shed low-priority work, or spill to durable storage. The choice depends on whether loss is acceptable.
Why it matters in production
Without a bound, latency and memory grow together until the process or dependency fails. A queue only moves the overload unless its capacity and policy are explicit.
Backpressure makes overload visible to callers and operators, allowing the system to degrade deliberately instead of failing unpredictably.
How it works
Bounded queue
A bounded queue has a full condition that must map to a policy: block, drop, reject, shed low-priority work, or spill to durable storage. The choice depends on whether loss is acceptable.
Flow across boundaries
Propagate demand or admission signals through streams, workers, HTTP limits, and broker consumers. An unbounded buffer between two bounded stages recreates the problem.
Fairness and priority
Separate tenants or priority classes when one producer can monopolise capacity. Measure queue age and service time, not only queue length.
Thickening: bounded buffers
Treat bounded buffers as an explicit budget. Record admission, queueing, and rejection separately so a full compartment cannot look like a healthy dependency.
Thickening: queue age and producer throttling
Exercise queue age and producer throttling under a deadline and a failed dependency. The useful signal is whether unrelated work keeps its capacity and whether recovery avoids a retry surge.
Implementing it
Build a bounded worker pool and test producers faster than consumers.
Choose a rejection response and retry contract for a full queue.
Add queue age, depth, processing rate, and dropped-work metrics.
Use a two-sided test for this boundary: drive the normal path and the failure path concurrently, then inspect the state that survives the race. For backpressure, the useful assertion is the invariant after recovery, not merely a successful response from one caller.
Document the limit and the signal that tells an operator to change it. A production review of backpressure should name the protected resource, the caller deadline, the expected overload decision, and the evidence that would distinguish a local bug from downstream saturation.
A focused review of backpressure should separate the mechanism from its policy. Reproduce one normal request, one boundary case, and one concurrent failure; record the state transition, the resource consumed, and the signal an operator would see. Then state what the caller is allowed to retry and what must be reconciled manually. This makes backpressure testable in a repository rather than a vocabulary answer.
Interview questions and how to answer them
What should happen when a queue is full?
Choose explicitly: block a producer, reject with a retry signal, drop a lower-priority item, or persist to durable storage. The right policy follows the value and loss tolerance of the work.
How is backpressure different from rate limiting?
Rate limiting controls admission against a policy over time; backpressure reflects current consumer capacity and propagates saturation through the pipeline. They often work together.
How do you size a queue?
Use a latency or burst budget, service rate, and memory limit. A queue should absorb a known burst, not hide unlimited overload.
What does a full buffer mean?
The consumer cannot keep up. Slow or reject the producer, persist work, or shed optional traffic explicitly; allowing memory to grow is not backpressure.
How do you choose a buffer bound?
Derive it from the latency budget and recovery rate, then test a sustained rate mismatch. Queue age should predict when the caller will miss its deadline.
Answers that lose the round
- Increasing queue size until memory fails.
- Returning 202 for work that was silently dropped.
- Retrying rejected work immediately without jitter or a deadline.
- Measuring only CPU while queue age grows.
- Raising the limit without measuring the protected resource.
- Letting a fallback path bypass the same bound.
FAQ
Can a message broker provide backpressure?
It can bound consumer concurrency and expose lag, but producers and consumers still need a policy for retention, priority, and overload.
Should an HTTP server block when workers are full?
Only within a bounded deadline. After that, reject or shed work so request threads and connections do not become another unbounded queue.
What metric signals trouble first?
Queue age often reflects user impact earlier than depth. Pair it with admission rejects, processing rate, and dependency latency.