Distributed systems
A webhook was delivered twice and the customer was charged twice
Written and reviewed by Sahil Srivastav
# access log: the provider retried because our first response was slow
10.4.2.11 - [14/Sep/2026:09:41:02 +0000] "POST /webhooks/payments HTTP/1.1" 499 0 "-" 11.004
10.4.2.11 - [14/Sep/2026:09:41:22 +0000] "POST /webhooks/payments HTTP/1.1" 200 31 "-" 0.214
# both deliveries carried the same provider event id
evt_3PqLk9aBcD type=payment.captured amount=29900 attempt=1
evt_3PqLk9aBcD type=payment.captured amount=29900 attempt=2
# and both were applied to the ledger
ledger_entries: 2 rows for payment pay_8sH2 (2 x 299.00 credited)What this error actually means
Webhook delivery is at-least-once, and the retry policy is the provider’s, not yours. Any delivery where the provider does not see a timely 2xx is retried — including the case where your handler completed the work perfectly and the response was lost, slow, or dropped by a proxy. The provider cannot distinguish "not processed" from "processed, acknowledgement lost", so it retries.
This means a slow handler is an active cause of duplicates rather than merely a performance concern. Most providers time out delivery in the low tens of seconds. A handler that does the payment reconciliation, the invoice write, the email, and the analytics call inline will sit near that limit under load, and every delivery that crosses it is retried while the first execution is still running — producing two concurrent executions of the same event, not two sequential ones.
The financial damage comes from a second property: most webhook side effects are not naturally idempotent. Appending a ledger entry, incrementing a balance, extending a subscription period, or issuing a refund are all "do it again and it happens again" operations. The provider gave you a stable event identifier precisely so you can refuse the second execution, and ignoring it is what converts an ordinary retry into a double charge.
Ordering is the companion problem. Retries interleave, so `payment.captured` can arrive after `charge.refunded`, and a subscription’s `updated` events can arrive out of sequence. Deduplicating by event id fixes repeats but not inversions; the durable answer is to key the effect on the business object and apply it only when the incoming state is newer than the stored state.
Causes, most common first
- 1The handler acknowledges after doing all the work, and the work is slow. Every operation inside the request — database writes, an email, a downstream API call — adds to the window in which the provider will give up and retry. The longer the handler, the higher the duplicate rate, and the duplicates arrive concurrently with the original.
- 2The event id is logged but not used as a uniqueness key. Extremely common. The handler records `evt_...` in a log line or an events table with no constraint, then proceeds unconditionally. The evidence of the duplicate is captured and the duplicate still happens.
- 3Deduplication by event id, but the effect applied before the claim commits. The handler inserts into a `webhook_events` table and then performs the effect in a separate transaction. A crash between the two loses the work; a concurrent retry that reads before the insert commits performs the effect twice. The claim and the effect must commit together.
- 4A proxy or gateway timeout in front of the handler. The application finished in 12 seconds; the load balancer cut the connection at 10. Your logs show success, the provider records a failure, and the retry arrives. This is the case where nobody believes the handler is at fault because, locally, it is not.
- 5Two environments subscribed to the same endpoint. A staging deployment left pointing at production webhook secrets, or a blue/green pair both receiving deliveries. Each processes the event once, correctly, and the business object is updated twice. Deduplication scoped per instance will not catch it; a shared store will.
When you see it
- Two ledger rows or two fulfilments for a single provider event id, seconds to minutes apart
- Access logs show 499, 502, or a long-latency 200 immediately before the duplicate
- Duplicates concentrate at traffic peaks and during downstream slowdowns
- The provider dashboard shows the delivery with attempt count greater than one and a successful final attempt
- A refund and its capture appear to have been applied in the wrong order
- Reconciliation against the provider’s own report finds your totals higher than theirs
How to diagnose it
Step 1
Read the provider’s delivery log for the event
Start here rather than in your own logs. The provider tells you how many attempts it made and what it saw for each — a slow 200, a 499, a 5xx. That distinguishes "we were slow" from "we errored" from "we were never reached", and each has a different fix.
stripe events resolve evt_3PqLk9aBcDStep 2
Measure webhook handler latency at the percentile that matters
The mean is irrelevant; the tail is what crosses the provider’s timeout. If p99 is anywhere near the provider’s delivery timeout, duplicates are not an anomaly, they are the expected steady state.
awk '$7 ~ /webhooks/ {print $NF}' access.log | sort -n | awk '{a[NR]=$1} END {print "p50", a[int(NR*0.5)], "p99", a[int(NR*0.99)], "max", a[NR]}'Step 3
Find the duplicated effects, not the duplicated deliveries
A duplicate delivery that was correctly ignored is harmless. Query for the business consequence — two entries against one payment — so you are measuring damage rather than noise.
SELECT payment_id, count(*), sum(amount_minor)
FROM ledger_entries
WHERE created_at > now() - interval '7 days'
GROUP BY payment_id HAVING count(*) > 1;Step 4
Check for a concurrent second execution rather than a sequential one
Log the request id and the event id on entry and exit. Overlapping intervals for the same event id prove concurrency, which rules out any fix based on checking first and writing later.
The fix
Acknowledge fast and process asynchronously. Validate the signature, persist the raw event with its provider event id under a unique constraint, return 200, and let a worker do the rest. This collapses handler latency to a single insert and removes the timeout-driven retries that cause most duplicates in the first place.
Make the provider event id the primary key of the stored event, and treat a conflict as success. `INSERT ... ON CONFLICT DO NOTHING` followed by a check of the affected-row count is atomic, so of two concurrent retries exactly one proceeds. Never implement this as a `SELECT` followed by an `INSERT`.
Apply the effect in the same transaction as the claim, or make the worker itself idempotent on the business key. If the claim commits and the process dies before the effect, the event is marked handled and the work is lost — a silent under-count that is harder to find than a duplicate.
Guard the business object against out-of-order events with a version or state condition, not just against repeats: extend a subscription period only if the new period end is later than the stored one, and apply a capture only if the payment is not already captured or refunded.
Verify the signature and reject unsigned deliveries before any of this. Deduplication logic keyed on an attacker-supplied event id is a replay vector, and the signature check is what makes the id trustworthy.
// Slow, unconditional, and duplicated on every provider retry.
@PostMapping("/webhooks/payments")
public ResponseEntity<Void> handle(@RequestBody String body) {
var evt = verifyAndParse(body);
ledger.credit(evt.paymentId(), evt.amountMinor()); // appends a row
invoices.markPaid(evt.paymentId());
mailer.sendReceipt(evt.customerId()); // ~2s on a good day
return ResponseEntity.ok().build();
}
// Claim atomically, acknowledge immediately, do the work once, out of band.
@PostMapping("/webhooks/payments")
public ResponseEntity<Void> handle(@RequestBody String body,
@RequestHeader("Stripe-Signature") String sig) {
var evt = verifyAndParse(body, sig); // throws -> 400, no retry loop
int claimed = jdbc.update("""
INSERT INTO webhook_events (event_id, type, payload, received_at)
VALUES (?, ?, ?::jsonb, now())
ON CONFLICT (event_id) DO NOTHING
""", evt.id(), evt.type(), body);
if (claimed == 1) {
worker.enqueue(evt.id()); // same transaction as the claim
}
return ResponseEntity.ok().build(); // sub-50ms, always
}
-- and the effect itself refuses to run twice or backwards
UPDATE payments
SET status = "captured", captured_at = now()
WHERE id = $1
AND status = "authorized"; -- 0 rows => already captured or refundedHow to stop it coming back
- Alarm on webhook handler p99 latency against the provider’s delivery timeout; this is the leading indicator of duplicates
- Make the unique constraint on the provider event id a schema-level requirement for every webhook table, reviewed like any other invariant
- Reconcile daily against the provider’s settlement or event export and alert on divergence — duplicates and losses are both silent without it
- Keep webhook secrets per environment so a stale staging deployment cannot process production events
- Replay a recorded delivery twice in CI, including concurrently, and assert one effect
Practise this failure in a real repository
The Gronex challenge hands you a payment ledger that double-credits on webhook redelivery and drops entries when the worker dies between the claim and the effect. The tests check the ledger balances against the event stream, so only a fix that makes claim and effect atomic passes.
FAQ
Why did the provider retry when my logs show a 200?
Because something between you and the provider did not deliver that 200 — most often a gateway or load-balancer timeout shorter than your handler, sometimes a client disconnect logged as 499. The provider records what it received, so trust its delivery log over your application log for this question.
Is deduplicating on the event id enough?
It stops repeats of the same event. It does not stop two *different* events from being applied in the wrong order, and it does not stop a retry of one event racing a fresh delivery of another. Pair the event-id claim with a state condition on the business object.
Should I return a non-2xx so the provider retries later?
Only for genuinely transient failures where you want redelivery. Returning 5xx for a permanently invalid payload puts the event into the provider’s retry schedule for days and can get your endpoint disabled. Return 400 for unprocessable input, 200 once the event is durably stored, and 5xx only when a retry can plausibly succeed.
How do I clean up a double charge that already happened?
Reconcile against the provider first to establish what actually happened on their side, then write a compensating ledger entry rather than deleting rows — a financial log should be append-only so the correction is auditable. Fix the constraint before the backfill, or the backfill will race new deliveries.
Related
Other errors engineers hit next to this one
- Undici / fetch connections remain occupied
- ERR_UNHANDLED_REJECTION in a worker thread
- command not found in a script that works interactively
- Permission denied when executing a script
- bad interpreter: No such file or directory with CRLF
- Argument list too long
- Too many open files
- Out of memory: Killed process