Distributed systems
The same message is processed twice — at-least-once delivery and idempotency
Written and reviewed by Sahil Srivastav
2026-09-14T11:02:17.884Z INFO OrderConsumer processed order=ORD-84213 partition=3 offset=91746 attempt=1
2026-09-14T11:02:47.113Z WARN OrderConsumer partitions revoked mid-batch: orders-3
2026-09-14T11:02:49.002Z INFO OrderConsumer processed order=ORD-84213 partition=3 offset=91746 attempt=1
2026-09-14T11:02:49.031Z ERROR OrderConsumer org.postgresql.util.PSQLException: ERROR: duplicate key value violates unique constraint "shipments_order_id_key"
Detail: Key (order_id)=(ORD-84213) already exists.What this error actually means
There is no exception for this one, which is precisely why it survives in production for months. The system works; it just occasionally does the same thing twice, and the evidence is a duplicate row, a second email, or a customer charged twice — discovered by a human long after the log lines have rotated away.
Redelivery is not a fault. Every broker in mainstream use — Kafka, SQS, RabbitMQ, Pub/Sub — delivers at least once, because the alternative is losing messages. The acknowledgement is a second network round trip after the work, and any failure in that window (a crash, a rebalance, a lost ack, a visibility timeout expiring while you were still working) leaves the broker with no evidence that the work happened. Its only safe move is to deliver again.
So the interesting question is never "why was it delivered twice" but "what in my handler is not safe to run twice". Handlers divide cleanly. A `SET status = "shipped"` is naturally idempotent and can run a thousand times. A `balance = balance - 10`, an `INSERT` without a unique constraint, a `POST /charges`, or an email send are not, and each needs an explicit key that lets the second execution recognise the first one.
The subtlety that catches careful teams: two deliveries can be *concurrent*, not sequential. During a rebalance or a visibility-timeout expiry the original worker is often still running. Any deduplication built as "select, then decide, then write" has a race between the select and the write, and under exactly the conditions that cause redelivery, both workers will pass the check.
Causes, most common first
- 1The acknowledgement was lost, not the work. The handler finished, the commit or ack did not reach the broker — a rebalance, a pod termination, an exhausted connection. From the broker’s position the message was never processed. This is the single most common origin and it is unavoidable by design.
- 2The visibility timeout or lease expired mid-processing. SQS returns the message to the queue after `VisibilityTimeout`; Kafka reassigns the partition after `max.poll.interval.ms`. Both fire while your worker is still running, so you get two workers on the same message at the same time — the concurrent case, not the sequential one.
- 3A retry after an ambiguous downstream result. The handler called a payment or shipping API, the call timed out, and it retried. The first call may well have succeeded server-side. Nothing in the message was redelivered — the duplication is inside your own retry, which makes it easy to miss when you go looking at the broker.
- 4Deduplication implemented as check-then-write. `if (!repo.exists(id)) repo.save(...)` is not deduplication under concurrency. Two workers execute the check before either writes. It reduces the duplicate rate enough to look fixed in testing and leaves the failure intact in production.
- 5A producer that publishes twice. The upstream service retried a publish after a timeout with no producer idempotence enabled, so two physically distinct messages carry the same business event. No consumer-side offset logic can help here — only a business key can.
When you see it
- Two identical business records minutes apart, with different message receipt ids
- Support tickets about duplicate emails or notifications that engineers cannot reproduce
- Ledger or inventory totals that drift from the sum of their source events
- Unique-constraint violations appearing in consumer logs and being caught and ignored
- Duplicates cluster around deploys, rebalances, and downstream timeouts rather than being uniform
How to diagnose it
Step 1
Prove the duplicate is real and measure it
Count duplicate business keys directly. This tells you the rate, and the time gap between copies tells you the mechanism: milliseconds means concurrent delivery, tens of seconds means a lease expiry or a rebalance, minutes means a retry after an ambiguous call.
SELECT order_id, count(*), max(created_at) - min(created_at) AS gap
FROM shipments
GROUP BY order_id HAVING count(*) > 1
ORDER BY 2 DESC LIMIT 20;Step 2
Decide whether the broker or your retry produced it
Log the broker-side identity — `topic-partition-offset` for Kafka, `messageId` and `ApproximateReceiveCount` for SQS — on every handler entry. Same identity twice means redelivery; different identities with the same business key means a double publish; one identity with an internal retry means your own client.
aws sqs receive-message --queue-url "$Q" --attribute-names ApproximateReceiveCount --max-number-of-messages 1Step 3
Audit each side effect for natural idempotency
Walk the handler and classify every write and every outbound call: idempotent as written, idempotent with a key, or not idempotent at all. The third category is your whole problem and it is usually two or three lines out of two hundred.
Step 4
Reproduce it deliberately in a test
Invoke the handler twice with the identical payload, and then invoke it twice concurrently from two threads. Most implementations pass the first test and fail the second, which is exactly the distinction that matters in production.
The fix
Give every message a stable business identity chosen by the producer — an order id, an event id, a payment intent id — and make it part of the payload. Broker-assigned identifiers are not enough, because a double publish produces two of them for the same event.
Enforce uniqueness in the database rather than in application logic. A `UNIQUE` constraint on the business key, combined with `INSERT ... ON CONFLICT DO NOTHING`, is atomic and correct under concurrency; a preliminary `SELECT` is neither. Treat the conflict as success and acknowledge the message.
Where the effect is an update rather than an insert, make the write conditional on state so the second execution is a no-op: `UPDATE orders SET status = "shipped" WHERE id = $1 AND status = "packed"`. Check the affected-row count to know whether you were first.
For external calls, pass an idempotency key derived from the business identity and let the provider deduplicate. Record your own outcome for that key in the same transaction as the rest of the work, so a replay finds the recorded result instead of calling again.
Keep the deduplication record and the business effect in one transaction. Writing the effect and then marking the message as seen leaves a window where a crash between them produces a duplicate on replay — which is the same class of bug as the dual write this whole family of failures comes from.
// Not deduplication: both workers pass the check before either inserts.
if (shipmentRepo.findByOrderId(evt.orderId()).isEmpty()) {
shipmentRepo.save(new Shipment(evt.orderId()));
carrier.dispatch(evt.orderId());
}
// Atomic claim first; only the winner performs the effect.
@Transactional
public void handle(OrderPacked evt) {
int claimed = jdbc.update("""
INSERT INTO processed_events (event_id, consumer, processed_at)
VALUES (?, "shipment-consumer", now())
ON CONFLICT (event_id, consumer) DO NOTHING
""", evt.eventId());
if (claimed == 0) return; // a previous delivery already did this
shipmentRepo.save(new Shipment(evt.orderId()));
carrier.dispatch(evt.orderId(), evt.eventId()); // key passed downstream too
}
-- the constraint that makes the claim atomic rather than advisory
CREATE TABLE processed_events (
event_id text NOT NULL,
consumer text NOT NULL,
processed_at timestamptz NOT NULL,
PRIMARY KEY (event_id, consumer)
);How to stop it coming back
- Write down the delivery semantics of every consumer in the service README: at-least-once with an idempotency key, or at-least-once with a naturally idempotent effect. There is no third option that is safe
- Add a contract test that delivers each message type twice, concurrently, and asserts the end state — this is cheap and catches regressions when a handler grows a new side effect
- Prune the deduplication table on a schedule with a retention window longer than your maximum redelivery delay, so it does not become the next incident
- Make producers idempotent where the broker supports it (`enable.idempotence=true` for Kafka) so double publishes stop at the source
- Alarm on duplicate business keys as a data-quality metric; the failure is silent otherwise and the log evidence expires before the ticket arrives
Practise this failure in a real repository
The Gronex repository delivers the same events twice — sequentially and concurrently — and the test suite asserts the resulting state rather than the code. A check-then-write fix fails the concurrent case, which is the distinction most candidates discover only here.
FAQ
Can I configure the broker for exactly-once delivery instead?
Not across a network boundary. What Kafka calls exactly-once semantics is an atomic read-process-write inside Kafka itself; the moment your handler touches an external database or API, delivery is at-least-once again and idempotency is your responsibility. Effectively-once processing on top of at-least-once delivery is the achievable goal.
Is a Redis SETNX good enough for deduplication?
It is fast and it is atomic, but it is a separate failure domain from your data. If the Redis write succeeds and the database write fails, the replay is suppressed and the work is lost — a worse outcome than a duplicate. Keep the dedup record in the same transaction as the effect, and use Redis only as an optimisation in front of it.
How long should I keep processed-event records?
Longer than the worst realistic redelivery gap: your queue retention for a DLQ replay, or your backlog replay window, whichever is larger. Days rather than minutes. Size it deliberately and index the cleanup, because this table grows at the rate of your traffic.
What if the side effect is an email and there is no key support?
Record the send in your own database inside the same transaction as the claim, and send after the commit. You then get at-most-once-looking behaviour with a small window where a crash after commit loses one email — which is almost always preferable to sending it twice. Make that choice explicitly rather than by accident.
Related
Other errors engineers hit next to this one
- celery.exceptions.WorkerLostError: Worker exited prematurely
- requests.exceptions.ReadTimeout: HTTPSConnectionPool read timed out
- UnicodeDecodeError: 'utf-8' codec can't decode byte
- AssertionError: daemonic processes are not allowed to have children
- [CRITICAL] WORKER TIMEOUT (pid:1234)
- Mutable default argument retains state across calls
- FATAL ERROR: Reached heap limit Allocation failed
- Unhandled promise rejection crashes the process