API design

API rate limiting: interview questions and practical design

API rate limiting bounds request or resource consumption over time so one caller cannot exhaust a shared service budget.

Written and reviewed by Sahil Srivastav

API designBackend engineeringInterview guide

What it actually is

API rate limiting bounds request or resource consumption over time so one caller cannot exhaust a shared service budget.

A limit protects availability and creates a fair allocation when traffic exceeds capacity or a client misbehaves.

The useful interview answer is precise about the boundary: A bucket holds tokens up to a burst capacity and refills at a rate. A request consumes tokens atomically; the next-available time can be calculated from the deficit instead of making clients guess.

Why it matters in production

A limit protects availability and creates a fair allocation when traffic exceeds capacity or a client misbehaves.

The limit must match the resource being protected: request count is not enough when requests have very different cost.

How it works

Token bucket

A bucket holds tokens up to a burst capacity and refills at a rate. A request consumes tokens atomically; the next-available time can be calculated from the deficit instead of making clients guess.

Identity and scope

Choose whether the key is an account, API key, IP, route, or tenant, and combine dimensions when one limit cannot protect the dependency.

Distributed accounting

Multiple instances need a shared or partitioned counter with an explicit consistency trade-off. A local counter is cheap but allows a cluster-wide burst on every replica.

Implementing it

Specify a limit, burst, identity, and response contract for one endpoint.

Test concurrent requests at the boundary and inspect `Retry-After` behaviour.

Model a dependency-specific budget separately from the public request limit.

Interview questions and how to answer them

Token bucket or leaky bucket?

Token bucket permits controlled bursts while bounding the long-term rate; a leaky bucket smooths output. Choose from the traffic shape and downstream tolerance.

Where should rate limiting happen?

At the edge for abusive traffic and at the resource owner for expensive operations. Layered limits are useful when those budgets differ.

How should a client respond to 429?

Respect `Retry-After` when present, apply bounded backoff with jitter, and avoid retrying permanently rejected requests.

Answers that lose the round

  • Using one global lock or counter for every endpoint.
  • Trusting an IP as the only identity behind proxies or NAT.
  • Returning 429 without telling the client when retrying may work.
  • Treating rate limiting as a replacement for capacity planning.

Practise in a real repository

Explaining a concept and enforcing it in code are different skills, and machine coding rounds test the second. Gronex ships broken backend repositories whose test suites assert the invariant rather than the happy path.

FAQ

Does rate limiting guarantee availability?

No. It reduces one class of overload; dependency failure, hot keys, and valid aggregate demand still need capacity and isolation controls.

Can Redis implement a limiter?

Yes, if the counter and expiry operation is atomic and its failure mode is explicit. A fail-open choice protects availability but weakens the budget.

Should limits be per user or per IP?

Use the identity that reflects responsibility and abuse risk, often a combination of authenticated tenant and network-level guard.

More backend concepts