Distributed systems

Exactly-once claims versus effectively-once in practice

Written and reviewed by Sahil Srivastav

Distributed systemsDelivery semanticsIdempotency
Transaction committed, producer timed out before acknowledgement; retry produced a duplicate charge

What this error actually means

Exactly-once is a scoped protocol property. A Kafka transaction can atomically commit consumed offsets and produced Kafka records, but it cannot atomically charge a card, send an email, or commit to an unrelated database.

A timeout after a commit leaves the caller unable to distinguish “did not commit” from “committed but acknowledgement was lost”. Retrying is necessary for availability and creates a duplicate unless the destination deduplicates it.

Effectively-once means repeated attempts converge to one business result. That is achieved at the side-effect boundary with an idempotency key, uniqueness constraint, or transactional inbox—not by a label on the transport.

Causes, most common first

  1. 1Guarantee does not cross systems. The transaction coordinator cannot roll back an external API.
  2. 2Unknown outcome after timeout. The operation may have committed before the response was lost.
  3. 3No destination deduplication. Every delivery is treated as new.

When you see it

  • Broker metrics report successful transactions while external duplicates exist
  • A producer timeout followed by retry creates two business records
  • Consumers replay after a rebalance
  • Tests pass until the acknowledgement is dropped

How to diagnose it

Step 1

Draw the commit boundary

List each resource and mark which transaction actually includes it.

Step 2

Inject lost acknowledgements

Drop the response after the destination commits and observe retry behaviour.

tc qdisc add dev eth0 root netem loss 10%

Step 3

Find duplicate business keys

Query destination records by request or event key.

The fix

Use a stable idempotency key derived from the business operation.

Store the key and result under a unique constraint in the destination transaction.

Use an inbox or outbox when database state and messages must move together.

Make consumers safe under replay and commit offsets only after durable handling.

Document the exact scope of every “exactly once” claim.

INSERT INTO payments(idempotency_key, amount) VALUES (:key, :amount)
ON CONFLICT (idempotency_key) DO NOTHING

How to stop it coming back

  • Test duplicate and lost-ack paths
  • Track duplicate suppression
  • Keep keys for the replay horizon
  • Avoid random keys generated on every retry
  • Review cross-system transaction boundaries

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Does Kafka exactly-once prevent duplicate emails?

No. Email is outside Kafka’s transaction and needs its own idempotency or reconciliation.

Is at-least-once acceptable?

Yes, when side effects are idempotent and the business result converges.

Can a unique key replace a transaction?

It prevents duplicate rows, but the transaction must still atomically record the key and result.

Related

Other errors engineers hit next to this one

Full error and symptom index →