API design
Bulk endpoint design: interview questions and practical design
Bulk endpoint design handles many logical operations while preserving bounded work, useful result reporting, and safe retry semantics for each item.
Written and reviewed by Sahil Srivastav
What it actually is
Bulk endpoint design handles many logical operations while preserving bounded work, useful result reporting, and safe retry semantics for each item.
A loop of single requests multiplies network and transaction overhead, while one unbounded bulk request can monopolise a service.
The useful interview answer is precise about the boundary: Set item and payload limits, process in chunks, and return a job resource for work that exceeds the synchronous budget. Limits should protect the database and downstream quotas, not just HTTP body size.
Why it matters in production
A loop of single requests multiplies network and transaction overhead, while one unbounded bulk request can monopolise a service.
Partial success is normal in bulk work, so the contract must identify each item and make retries safe without repeating successful effects.
How it works
Bounded batches
Set item and payload limits, process in chunks, and return a job resource for work that exceeds the synchronous budget. Limits should protect the database and downstream quotas, not just HTTP body size.
Per-item outcome
Return a stable item identifier, status, error code, and result or location. Do not make a client infer which items succeeded from an aggregate 200.
Idempotent replay
Give the batch and each item a key or use a natural operation identity. A retry after a timeout must recognise completed items and avoid duplicating side effects.
Implementing it
Design synchronous bulk create with a maximum size and per-item result.
Add a partial failure and retry after the response is lost.
Move an oversized request to an asynchronous job with status polling or a webhook.
Interview questions and how to answer them
When should bulk work be asynchronous?
When processing time or fan-out can exceed the request deadline, or when progress and retry need a durable job state. Return a trackable job rather than holding a connection.
How do you preserve atomicity?
State whether atomicity is per item, per chunk, or for the whole batch. A single transaction may be correct for a small bounded batch but dangerous at scale.
How should clients retry partial results?
Retry only failed or unknown items with stable identities, and let the server deduplicate items already committed.
Answers that lose the round
- Wrapping a loop in one huge transaction by default.
- Returning only the count of failures with no item identity.
- Retrying the entire batch when some effects already committed.
- Accepting arbitrary batch sizes and discovering the limit through timeouts.
FAQ
Should a bulk endpoint always be faster?
It can reduce overhead, but it still competes for CPU, locks, and downstream capacity. Bounded work and measured throughput matter more than request count.
What status should a bulk response use?
Use a documented result model. 200 can report per-item outcomes; 207 may be appropriate when the API adopts a multi-status contract. Consistency is key.
How do I test bulk APIs?
Test empty and maximum batches, duplicate items, partial failure, timeout after commit, ordering if promised, and permission isolation per item.