Distributed systems

At-least-once vs exactly-once delivery: interview questions and how to answer them

Exactly-once delivery is unachievable over a network that can lose messages; exactly-once *effects* are achievable, by combining at-least-once delivery with deduplication or idempotent processing.

Written and reviewed by Sahil Srivastav

Distributed systemsMessagingCommonly misstated

What it actually is

Delivery semantics describe what a messaging system promises when things fail. At-most-once means a message may be lost but never duplicated — you acknowledge before processing. At-least-once means a message may be duplicated but never lost — you acknowledge after processing. Exactly-once would mean neither, and as a property of *delivery* over an unreliable channel it cannot be provided.

The impossibility is not a limitation of current engineering. It follows from the same reasoning as the Two Generals problem: the sender cannot distinguish “my message was lost” from “the reply was lost”, so it must either resend (risking a duplicate) or not resend (risking a loss). No finite protocol removes that choice, because any acknowledgement can itself be the message that is lost.

What is achievable is exactly-once *processing*, sometimes called effectively-once: the message may arrive many times, but the observable effect happens once. That is a property of the consumer plus its storage, not of the transport. Every production system that advertises exactly-once is doing this, and the interesting question is where it keeps the deduplication state.

Why it matters in production

Because the difference is a duplicate charge. A consumer that debits a wallet on each delivery and relies on the broker never redelivering will eventually debit twice — a consumer crash in the window between the side effect and the acknowledgement is enough, and that window cannot be closed by making the code faster.

It also decides where you spend engineering effort. If you accept at-least-once, you invest in idempotency keys, unique constraints, and guarded state transitions. If you believe the transport gives you exactly-once, you invest in nothing and discover the gap in production, usually during the incident that follows a broker rebalance.

And the semantics you need differ by message. Metrics samples are fine at-most-once — a dropped gauge reading costs nothing and a duplicate corrupts the aggregate. Payment instructions must be at-least-once with dedup. Treating all traffic with one policy over-engineers the cheap path and under-engineers the expensive one.

How it works

Where the duplicate comes from

Processing and acknowledging are two separate operations against two separate systems, so there is always an instant when one has happened and the other has not. Crash there and the broker’s redelivery timer fires: the work is done, the offset is unmoved, the message comes back. Consumer-group rebalances create the same window deliberately, which is why duplicates cluster around deploys and scaling events.

Dedup needs a key and a store

Deduplication is a unique constraint plus a retention window. The key must identify the logical event — a producer-assigned event id, or a natural key such as `(order_id, transition)` — and the store must be the same store the side effect writes to, so that claiming the key and applying the effect commit together. A Redis `SETNX` guarding a database write reintroduces the gap it was meant to close.

What Kafka’s exactly-once actually covers

Two mechanisms. The idempotent producer attaches a producer id and a per-partition sequence number, so a broker rejects a retried duplicate of the same batch — this removes producer-side duplicates, not consumer-side ones. Transactions then let a stream processor atomically commit its consumed offsets and its produced records inside Kafka, and readers set `isolation.level=read_committed` to skip aborted output. The guarantee is exactly-once for read-process-write *within Kafka*.

The sink is where it breaks

The moment the effect lands outside the transactional boundary — an HTTP call, an email, a row in another database — Kafka’s transaction cannot include it, so you are back to at-least-once at that edge. Frameworks bridge this with a two-phase-commit sink (pre-commit on checkpoint, commit on notification) or, far more commonly and more cheaply, with an idempotent write into the sink.

Effectively-once as a contract

The useful formulation for a design review: the transport guarantees at-least-once, the consumer guarantees idempotent application, and together they guarantee the effect occurs once. Stating it that way makes the ownership explicit — and makes it obvious that adding a second, non-idempotent side effect to the handler breaks the guarantee for the whole pipeline.

Implementing it

Acknowledge after the effect is durable, never before, unless the message is genuinely disposable. That choice alone is what makes the pipeline at-least-once rather than at-most-once.

Have producers stamp a stable event id at creation time and carry it through every hop. Deriving an id from the payload hash breaks when the payload legitimately repeats — two identical “heartbeat” events an hour apart are different events.

Put the dedup record and the business write in one transaction. If the effect is external, write intent locally and let a separate worker perform the call, which is the transactional outbox pattern; it converts an unsolvable atomicity problem into a solvable retry problem.

Bound the dedup window explicitly and size it against the worst redelivery horizon you allow — including a message replayed from a dead-letter queue days later. A window shorter than that silently degrades to at-least-once.

-- Consumer: claim the event and apply the effect in one transaction
BEGIN;

INSERT INTO processed_events (event_id, consumer, processed_at)
VALUES ($1, 'wallet-debiter', now())
ON CONFLICT (event_id, consumer) DO NOTHING;
-- 0 rows => already processed; ROLLBACK and ack the message.

UPDATE wallets SET balance_minor = balance_minor - $2
 WHERE id = $3 AND balance_minor >= $2;

COMMIT;
-- Only now acknowledge to the broker. A crash before the ack
-- redelivers the message, and the ON CONFLICT makes that harmless.

Interview questions and how to answer them

Can a message broker guarantee exactly-once delivery? Explain.

No. The sender cannot distinguish a lost message from a lost acknowledgement, so it must choose between resending and not resending — duplicates or loss. What systems can guarantee is exactly-once *effect*: at-least-once delivery plus an idempotent consumer, with the dedup state committed atomically with the side effect.

Then what does Kafka mean by exactly-once semantics?

Two things. The idempotent producer eliminates duplicate batches caused by producer retries, using a producer id and per-partition sequence numbers checked at the broker. Transactions let a processor commit consumed offsets and produced records atomically, with consumers reading at `read_committed`. Both are scoped to Kafka; a side effect in an external system is outside the transaction and still needs idempotency.

When is at-most-once the right choice?

When a duplicate is more harmful than a loss and the data is statistical rather than transactional: high-rate metrics, sampled traces, cache-warming hints, presence pings. A dropped sample barely moves an aggregate; a duplicated one skews it, and the cost of a dedup store per sample is not worth paying.

Your consumer crashes after charging a card but before acknowledging. What happens, and how do you make it safe?

The broker redelivers and the naive handler charges again. Safe version: before charging, claim the event id in the same database the charge record lives in, under a unique constraint; if the claim fails, the charge already happened and the handler acknowledges without repeating it. If the charge is at a third party, send its own idempotency key so the gateway collapses the retry.

How do you choose the deduplication key?

It must identify the logical operation, be assigned by the producer before the first send, and be stable across retries. A broker-assigned message id fails the last test after a redelivery through a different path; a payload hash fails when identical payloads are genuinely distinct events. A producer-generated UUID stored with the event, or a composite business key such as `(order_id, status_transition)`, works.

Is at-least-once enough on its own?

Only if every effect in the handler is naturally idempotent — absolute assignment, a guarded state transition, or a set insert. The moment the handler appends a ledger row, increments a counter, or sends a notification, at-least-once without dedup is a correctness bug rather than a rough edge.

Answers that lose the round

  • Claiming a broker provides exactly-once delivery — no broker does, and the ones that advertise it are describing exactly-once processing within their own boundary
  • Acknowledging before processing to “avoid duplicates”, which converts a duplicate problem into silent message loss
  • Deduplicating in a cache while writing the effect to a database, so a crash between the two leaves the key set and the work undone
  • Relying on Kafka transactions for a side effect that leaves Kafka, such as an outbound HTTP call or an email
  • Using a payload hash as the dedup key, which collapses legitimately identical events
  • Assuming duplicates are rare enough to ignore; rebalances and deploys make them routine, not exceptional
  • Forgetting that a retry from a dead-letter queue can arrive long after the dedup window has expired

Practise delivery semantics in a real repository

This Gronex repository is a webhook consumer that the test suite deliberately abuses: the same event is redelivered, events arrive out of order, and a crash is simulated between the side effect and the acknowledgement. Final state must be correct, so a dedup check with a race window or a cache-based guard does not pass.

FAQ

What is the difference between exactly-once delivery and exactly-once processing?

Delivery is about the transport putting the message in front of the consumer exactly one time — impossible. Processing is about the effect occurring one time regardless of how many deliveries happen — achievable, and the thing you actually want. Interviewers listen for whether a candidate keeps the two apart.

Does SQS FIFO give exactly-once?

SQS FIFO offers exactly-once *processing* within a five-minute deduplication window using a message deduplication id, and strict ordering within a message group. Outside that window a resend is a new message, so a consumer still needs its own idempotency for anything with a longer retry horizon.

Do I need both an outbox and consumer-side dedup?

Usually yes, because they cover different halves. The outbox makes “write the row and publish the event” atomic on the sending side, which guarantees at-least-once publication. Consumer dedup makes at-least-once safe to receive. Neither substitutes for the other.

Related

More backend concepts