API design
API error handling: interview questions and practical design
API error handling turns validation, domain rejection, conflicts, and infrastructure failures into stable signals without leaking implementation details.
Written and reviewed by Sahil Srivastav
What it actually is
API error handling turns validation, domain rejection, conflicts, and infrastructure failures into stable signals without leaking implementation details.
Clients need to know whether to fix input, refresh state, wait, or stop retrying.
The useful interview answer is precise about the boundary: Separate malformed input, failed business rule, authentication, authorisation, conflict, throttling, and transient server failure. Each category has different remediation and retry semantics.
Why it matters in production
Clients need to know whether to fix input, refresh state, wait, or stop retrying.
Centralised error mapping prevents one endpoint from returning a stack trace, another a string, and a third a misleading success response.
How it works
Error taxonomy
Separate malformed input, failed business rule, authentication, authorisation, conflict, throttling, and transient server failure. Each category has different remediation and retry semantics.
Stable envelope
Return a machine code, safe message, optional field details, and correlation id. Keep internal exception classes out of the public contract.
Boundary mapping
Catch known domain errors near the transport boundary and let unexpected errors reach central logging and a generic 5xx response. Do not catch broadly inside every method.
Implementing it
Create an error catalogue and map each code to client action.
Test redaction of secrets and identifiers from errors and logs.
Simulate a downstream timeout and assert a bounded, retryable response.
Interview questions and how to answer them
Where should errors be translated?
Keep domain errors meaningful internally and translate them at the API boundary into the stable wire contract. A central handler provides consistency.
How do you avoid leaking information?
Use allow-listed details, redact tokens and personal data, and give support a correlation id rather than a stack trace.
Should error codes be globally unique?
They should be stable within the API contract and documented. Scope them by domain if that keeps ownership clear, but do not reuse a code for different client actions.
Answers that lose the round
- Returning exception messages directly to clients.
- Logging the same error at every layer, multiplying noise while still losing context.
- Treating all 4xx responses as safe to retry.
- Using a single `UNKNOWN_ERROR` code for actionable domain failures.
FAQ
What is a good error response?
A predictable status, stable code, safe message, actionable details, and correlation id. The exact JSON shape should be consistent across endpoints.
Should validation happen in the controller?
Transport shape validation can happen there; domain rules belong in reusable application or domain code.
How do errors affect retries?
Classify them explicitly. Permanent input and authorisation errors should stop retries; transient capacity or dependency errors need bounded backoff and a deadline.