Performance
Circuit breaker: interview questions and practical design
A circuit breaker stops sending work to a failing dependency for a bounded period, giving the dependency and the caller a chance to recover instead of multiplying the outage.
Written and reviewed by Sahil Srivastav
What it actually is
A circuit breaker stops sending work to a failing dependency for a bounded period, giving the dependency and the caller a chance to recover instead of multiplying the outage.
Retries and concurrent callers can turn a slow dependency into thread, connection, and queue exhaustion.
The useful interview answer is precise about the boundary: Closed permits calls while tracking failures; open rejects quickly for a cool-down; half-open allows a limited probe before closing or reopening. Count meaningful failures, not every business rejection.
Why it matters in production
Retries and concurrent callers can turn a slow dependency into thread, connection, and queue exhaustion.
A breaker is an admission control mechanism, not a repair: it needs timeouts, fallback semantics, and a half-open probe policy.
How it works
Closed, open, half-open
Closed permits calls while tracking failures; open rejects quickly for a cool-down; half-open allows a limited probe before closing or reopening. Count meaningful failures, not every business rejection.
Timeout and budget
The breaker cannot react to a call that has no deadline. Set a dependency timeout below the caller budget and bound concurrent probes and in-flight work.
Fallback correctness
A fallback may serve stale data, queue work, or return an explicit unavailable response. It must not fabricate success or hide data that callers require to be fresh.
Thickening: half-open probes
Treat half-open probes as an explicit budget. Record admission, queueing, and rejection separately so a full compartment cannot look like a healthy dependency.
Thickening: state transition metrics
Exercise state transition metrics under a deadline and a failed dependency. The useful signal is whether unrelated work keeps its capacity and whether recovery avoids a retry surge.
Implementing it
Model a dependency that becomes slow before it becomes unavailable.
Test breaker state transitions with concurrent calls and a failed half-open probe.
Define metrics for rejection, dependency failures, open duration, and fallback use.
Use a two-sided test for this boundary: drive the normal path and the failure path concurrently, then inspect the state that survives the race. For circuit breaker, the useful assertion is the invariant after recovery, not merely a successful response from one caller.
Document the limit and the signal that tells an operator to change it. A production review of circuit breaker should name the protected resource, the caller deadline, the expected overload decision, and the evidence that would distinguish a local bug from downstream saturation.
A focused review of circuit breaker should separate the mechanism from its policy. Reproduce one normal request, one boundary case, and one concurrent failure; record the state transition, the resource consumed, and the signal an operator would see. Then state what the caller is allowed to retry and what must be reconciled manually. This makes circuit breaker testable in a repository rather than a vocabulary answer.
Interview questions and how to answer them
What opens a circuit?
A policy over meaningful failures, timeouts, or latency—not a single arbitrary exception. Use a rolling window and account for traffic volume so a quiet service does not flap on one sample.
What happens in half-open?
Allow a small, controlled number of probes. Success closes the breaker; failure reopens it for another bounded cool-down.
Is a circuit breaker enough for resilience?
No. Pair it with deadlines, bounded concurrency, sensible retries, bulkheads, and a fallback that preserves correct semantics.
What belongs in the half-open state?
A small, controlled probe set with normal timeouts. Close only after sustained evidence, otherwise reopen without allowing recovery traffic to become an outage.
How do you diagnose an open breaker?
Inspect transition counts, slow-call rate, dependency error class, and probe results. The breaker protects callers; it does not identify why the dependency failed.
Answers that lose the round
- Opening on all 4xx responses or domain rejections.
- Using a long timeout so the breaker reacts only after pools are exhausted.
- Allowing every instance to probe simultaneously.
- Returning cached or empty data without declaring its freshness.
- Raising the limit without measuring the protected resource.
- Letting a fallback path bypass the same bound.
FAQ
Should breakers be global or per dependency?
Usually per dependency and meaningful operation or tenant class. A global breaker can hide healthy paths behind one failing endpoint.
How long should the open period be?
Long enough to avoid hammering recovery, short enough to detect repair. Base it on dependency recovery and use jitter to avoid synchronized probes.
Can a breaker improve latency?
It improves tail latency during dependency failure by rejecting quickly, but healthy-path latency still depends on the dependency and timeout budget.