Data consistency
Transactional outbox: interview questions and how to answer them
A transactional outbox stores the intended message beside the business write, then publishes it asynchronously so a crash cannot lose the intent between two systems.
Written and reviewed by Sahil Srivastav
What it actually is
The service writes its domain row and an outbox row in one local transaction. A relay reads pending outbox rows and publishes them, marking them dispatched only after the broker acknowledges. A crash can produce a duplicate publish, so consumers must be idempotent.
The outbox solves database-to-broker atomicity, not global exactly-once delivery. Ordering, retention, relay concurrency, and poison messages remain design decisions.
Why it matters in production
Publishing after commit can lose the event if the process crashes first; publishing before commit can announce data that later rolls back. The outbox makes the durable intent unambiguous.
It is a practical foundation for sagas, search indexing, cache invalidation, and audit feeds.
How it works
One transaction
Business mutation and outbox insert commit together. The outbox payload should contain the event id, aggregate id, type, and schema version.
Relay claim
Workers claim rows with a lease or `SKIP LOCKED`, publish, then mark success. Lease expiry makes crashes recoverable but permits duplicates.
Ordering
Partition by aggregate or sequence events per aggregate. Global ordering is expensive and usually unnecessary.
Retention
Delete or archive only after the retry and replay window. Keep enough metadata to diagnose and rebuild downstream projections.
Detailed boundary
co-committing intent with domain state
Operational consequence
relay retries and deduplication
Implementing it
Use a unique event id and idempotent consumers. Treat the relay as at-least-once.
Alert on oldest pending age, publish failures, lease expiry, and table growth.
Version event schemas and make replay a supported operation, not an emergency SQL script.
Exercise relay crash windows: acknowledge the broker, terminate the worker before the database mark, and verify the duplicate is harmless. Measure oldest unpublished age separately from broker lag.
Document retention, replay scope, and poison-message handling so an operator can recover a stuck event without editing business rows.
Use a monotonic aggregate sequence when consumers need per-entity order, and keep the event payload self-contained enough for replay after the source schema changes.
BEGIN;
UPDATE orders SET status = 'PAID' WHERE id = $1;
INSERT INTO outbox(id, type, aggregate_id, payload)
VALUES ($2, 'OrderPaid', $1, $3);
COMMIT;Interview questions and how to answer them
What failure does the outbox close?
The crash window between a committed database change and its message publish, or between publish and recording that it was sent.
Can the relay publish duplicates?
Yes. A crash after broker acceptance and before marking the row sent causes a retry. Consumers must deduplicate.
How do you preserve order?
Use an aggregate sequence and partition or claim events in that order; do not assume concurrent relay workers preserve table order.
Why not use 2PC?
A broker may not participate, and 2PC can block. The outbox trades global atomicity for durable intent and idempotent eventual delivery.
What evidence would you inspect for transactional outbox?
Measure the boundary named in the design, compare it with the caller deadline and resource budget, and reproduce the contention or failure with more than one concurrent worker.
What is the tempting fix for this problem?
Changing a timeout, pool, or retry count alone usually moves the queue. First establish the invariant, then make the bounded mechanism and its failure outcome explicit.
Answers that lose the round
- Publishing directly after a database commit
- Marking dispatched before broker acknowledgement
- Assuming one relay attempt means one delivery
- Using a random event id on every retry
- Deleting rows before downstream replay is possible
- Putting remote calls inside the database transaction
- Treating the local mechanism as a complete production guarantee
- Changing the limit without measuring the resource it protects
FAQ
Does the outbox remove eventual consistency?
No. It makes the delay and retry durable; consumers still apply the event asynchronously.
Where should deduplication live?
With the consumer’s side effect, ideally in the same transaction as applying that effect.
Can an outbox grow forever?
No. Retention, archiving, and replay policy must be explicit and monitored.