Distributed systems

Idempotency: interview questions and how to answer them

An operation is idempotent when performing it more than once has the same effect as performing it once — the property that makes retrying safe.

Written and reviewed by Sahil Srivastav

Distributed systemsReliabilityAsked constantly

What it actually is

An operation is idempotent if applying it repeatedly leaves the system in the same state as applying it once. Note the definition is about *effect*, not about the response: returning a cached result is one way to be idempotent, and so is performing the work again in a way that converges on the same state.

The distinction that trips people up is between naturally idempotent and made idempotent. `SET balance = 100` is naturally idempotent because it is absolute. `balance = balance - 10` is not, because it is relative — applying it twice removes twenty. Most business operations are relative, so idempotency is something you engineer rather than something you have.

It is also worth being precise that idempotency is not the same as deduplication, though deduplication is the usual implementation. Deduplication remembers what it has seen; idempotency is the guarantee that duplicates do not change the outcome. You can achieve the guarantee without remembering anything if the operation is absolute or conditional.

Why it matters in production

Because the network cannot tell you whether a request succeeded. A client sends a payment request and the connection drops before the response arrives. The payment may have been captured or not — the client genuinely cannot know. Its only options are to retry (risking a double charge) or not retry (risking a lost payment). Idempotency is what makes the first option safe, and therefore what makes the system usable at all.

Every message broker and webhook sender in production use offers at-least-once delivery, not exactly-once. Kafka, SQS, Stripe webhooks, and payment gateway callbacks will all deliver the same message twice under the right failure conditions — a consumer crash after processing but before acknowledging is enough. So a consumer that is not idempotent is not merely fragile; it is incorrect by construction, and it will produce a duplicate charge eventually.

This is why the concept appears in nearly every backend interview at a payments, logistics, or booking company. It is the single cheapest way to find out whether a candidate has thought about partial failure.

How it works

The idempotency key

The client generates a unique key per logical operation — not per attempt — and sends it with every retry of that operation. The server stores it under a unique constraint. The key must be generated before the first attempt and reused across retries, which is the part clients most often get wrong: a key regenerated per attempt provides no protection at all.

Atomic claim, not check-then-act

The server must claim the key and do the work in one atomic step. Checking whether the key exists and then inserting leaves a race window where two concurrent retries both pass the check. `INSERT ... ON CONFLICT DO NOTHING RETURNING` claims and reports in a single statement; the unique index, not the check, is what enforces the guarantee.

Storing the response, not just the key

A retry should receive the original outcome, which means persisting the response alongside the key. Storing only the key lets you avoid duplicate work but forces you to return something vague on retry — and clients that cannot distinguish "already done, here is the result" from "unknown" will keep retrying.

Scope and expiry

A key is unique within a scope: per endpoint, per tenant, per account. Globally unique keys leak information across tenants and create false conflicts. Keys also need a retention window — long enough to outlast any client retry policy, short enough that the table does not grow forever. Twenty-four hours to seven days is typical.

Conditional updates as an alternative

Sometimes you do not need a key at all. `UPDATE inventory SET qty = qty - 1 WHERE id = $1 AND qty >= 1` is safe to repeat only if you also guard against reapplication — but `UPDATE orders SET status = 'PAID' WHERE id = $1 AND status = 'PENDING'` is naturally idempotent, because the second application matches no rows. State machines with guarded transitions get idempotency for free.

Implementing it

Require an idempotency key on every mutating endpoint a client can retry, and reject requests without one rather than making it optional — an optional safety property is not a safety property.

Claim the key in the same transaction as the work. If the work involves an external call, the key and the intent go into the database together and a separate worker performs the call; that is the transactional outbox pattern, and it exists because you cannot make a database write and a remote call atomic any other way.

Return the original response on a duplicate, with the same status code, so the client sees a successful retry rather than a conflict. Reserve `409` for a key reused with a *different* request body, which is a client bug worth surfacing loudly.

Make consumers idempotent at the message level too: store processed message ids, or derive a natural key from the event, and make the processing step a guarded state transition.

-- Claim and report in one statement: no race window
INSERT INTO payment_requests (idempotency_key, account_id, amount_minor, status)
VALUES ($1, $2, $3, 'PENDING')
ON CONFLICT (idempotency_key) DO NOTHING
RETURNING id;

-- 0 rows => this key was already claimed. Return the stored response:
SELECT response_status, response_body
  FROM payment_requests
 WHERE idempotency_key = $1;

-- Guarded transition: naturally idempotent, no key table needed
UPDATE orders SET status = 'PAID', paid_at = now()
 WHERE id = $1 AND status = 'PENDING';
-- 0 rows affected => already paid; that is success, not an error.

Interview questions and how to answer them

A client sends a payment request and the connection drops. What should it do?

Retry with the same idempotency key it used for the first attempt. The key is what lets the server distinguish "this is the same logical payment" from "this is a second payment", which the network cannot tell it. Without a key, neither retrying nor not retrying is safe — one risks a double charge and the other risks losing the payment silently.

Is a POST that creates a resource idempotent? How would you make it so?

Not by default — two identical POSTs create two resources, which is usually correct. To make it idempotent you need a client-supplied key stored under a unique constraint, claimed atomically with the creation, with the response persisted so a retry returns the original resource rather than creating a second one.

Your message broker guarantees at-least-once delivery. What does that require of your consumer?

That processing be idempotent, because duplicates are certain rather than unlikely — a crash between processing and acknowledging produces one. In practice: a unique constraint on the event id, or a guarded state transition that a second application cannot re-apply. "Exactly-once" end to end is achieved by an idempotent consumer, not by the transport.

What is wrong with checking whether the idempotency key exists and then inserting it?

There is a window between the check and the insert. Two concurrent retries of the same request both find no key, both proceed, and the work happens twice — which is precisely the scenario the key was added to prevent. The fix is one atomic statement, letting the unique index enforce the invariant.

How long should you retain idempotency keys, and what happens after that?

Longer than any client retry window — typically 24 hours to 7 days — then expire them, or the table grows without bound. After expiry a replayed request is treated as new, which is why the window must exceed the maximum retry horizon of every client, including ones retrying from a dead-letter queue days later.

Give an operation that is naturally idempotent and one that is not.

Absolute assignment is naturally idempotent: `SET status = 'PAID'`, or `PUT` of a full resource representation. Relative mutation is not: `balance = balance - 10`, appending to a list, incrementing a counter. The useful move in design is converting relative operations into guarded absolute ones — `WHERE status = 'PENDING'` makes the second application a no-op.

Answers that lose the round

  • Generating a new idempotency key on each retry attempt, which makes the whole mechanism a no-op
  • Check-then-insert on the key, which races under exactly the concurrent-retry conditions it is meant to handle
  • Storing the key but not the response, so retries get an ambiguous answer and keep retrying
  • Claiming "we use exactly-once delivery" — brokers give at-least-once, and effectively-once comes from idempotent consumers, not from the transport
  • Making the key globally unique instead of scoped, producing cross-tenant conflicts and information leaks
  • Treating a duplicate as an error and returning 409 when the correct answer is the original success response
  • Assuming `GET` and `PUT` are automatically safe: `PUT` is idempotent only if the handler actually sets absolute state

Practise idempotency in a real repository

Gronex ships this as a runnable repository: a webhook consumer facing duplicate and out-of-order deliveries. The tests redeliver events and replay them in the wrong order, asserting final state is correct — so a deduplication check with a race window does not pass.

FAQ

Is idempotency the same as exactly-once delivery?

No. Exactly-once delivery is not achievable over an unreliable network — the sender can never know whether a lost acknowledgement means "not delivered" or "delivered, ack lost". What is achievable is at-least-once delivery plus idempotent processing, which produces exactly-once *effects*. That is what systems claiming exactly-once are actually doing.

Should the client or the server generate the idempotency key?

The client, because only the client knows that two attempts are the same logical operation. A server-generated key cannot link a retry to the original request — the retry arrives as a fresh request with no way to identify it.

Do I need idempotency keys if my operations are already conditional?

Often not. A guarded state transition is naturally idempotent, so a second application matches no rows and changes nothing. Keys earn their cost when the operation genuinely appends or accrues — a ledger entry, a notification, an external charge.

How does this relate to the transactional outbox?

They solve adjacent halves of the same problem. Idempotency makes duplicate delivery harmless on the receiving side; the outbox makes a database write and a message publish atomic on the sending side. A reliable pipeline typically needs both — the outbox guarantees at-least-once send, idempotency makes at-least-once safe to consume.

Related

More backend concepts