Distributed systems
Exactly-once claims versus effectively-once in practice
Written and reviewed by Sahil Srivastav
Transaction committed, producer timed out before acknowledgement; retry produced a duplicate chargeWhat this error actually means
Exactly-once is a scoped protocol property. A Kafka transaction can atomically commit consumed offsets and produced Kafka records, but it cannot atomically charge a card, send an email, or commit to an unrelated database.
A timeout after a commit leaves the caller unable to distinguish “did not commit” from “committed but acknowledgement was lost”. Retrying is necessary for availability and creates a duplicate unless the destination deduplicates it.
Effectively-once means repeated attempts converge to one business result. That is achieved at the side-effect boundary with an idempotency key, uniqueness constraint, or transactional inbox—not by a label on the transport.
Causes, most common first
- 1Guarantee does not cross systems. The transaction coordinator cannot roll back an external API.
- 2Unknown outcome after timeout. The operation may have committed before the response was lost.
- 3No destination deduplication. Every delivery is treated as new.
When you see it
- Broker metrics report successful transactions while external duplicates exist
- A producer timeout followed by retry creates two business records
- Consumers replay after a rebalance
- Tests pass until the acknowledgement is dropped
How to diagnose it
Step 1
Draw the commit boundary
List each resource and mark which transaction actually includes it.
Step 2
Inject lost acknowledgements
Drop the response after the destination commits and observe retry behaviour.
tc qdisc add dev eth0 root netem loss 10%Step 3
Find duplicate business keys
Query destination records by request or event key.
The fix
Use a stable idempotency key derived from the business operation.
Store the key and result under a unique constraint in the destination transaction.
Use an inbox or outbox when database state and messages must move together.
Make consumers safe under replay and commit offsets only after durable handling.
Document the exact scope of every “exactly once” claim.
INSERT INTO payments(idempotency_key, amount) VALUES (:key, :amount)
ON CONFLICT (idempotency_key) DO NOTHINGHow to stop it coming back
- Test duplicate and lost-ack paths
- Track duplicate suppression
- Keep keys for the replay horizon
- Avoid random keys generated on every retry
- Review cross-system transaction boundaries
FAQ
Does Kafka exactly-once prevent duplicate emails?
No. Email is outside Kafka’s transaction and needs its own idempotency or reconciliation.
Is at-least-once acceptable?
Yes, when side effects are idempotent and the business result converges.
Can a unique key replace a transaction?
It prevents duplicate rows, but the transaction must still atomically record the key and result.
Related
Other errors engineers hit next to this one
- Task was destroyed but it is pending!
- Executing <Handle ...> took 2.418 seconds (blocked event loop)
- SettingWithCopyWarning: A value is trying to be set on a copy of a slice
- celery.exceptions.WorkerLostError: Worker exited prematurely
- requests.exceptions.ReadTimeout: HTTPSConnectionPool read timed out
- UnicodeDecodeError: 'utf-8' codec can't decode byte
- AssertionError: daemonic processes are not allowed to have children
- [CRITICAL] WORKER TIMEOUT (pid:1234)