Data consistency

Transactional outbox: interview questions and how to answer them

A transactional outbox stores the intended message beside the business write, then publishes it asynchronously so a crash cannot lose the intent between two systems.

Written and reviewed by Sahil Srivastav

Data consistencyBackend engineeringInterview preparation

What it actually is

The service writes its domain row and an outbox row in one local transaction. A relay reads pending outbox rows and publishes them, marking them dispatched only after the broker acknowledges. A crash can produce a duplicate publish, so consumers must be idempotent.

The outbox solves database-to-broker atomicity, not global exactly-once delivery. Ordering, retention, relay concurrency, and poison messages remain design decisions.

Why it matters in production

Publishing after commit can lose the event if the process crashes first; publishing before commit can announce data that later rolls back. The outbox makes the durable intent unambiguous.

It is a practical foundation for sagas, search indexing, cache invalidation, and audit feeds.

How it works

One transaction

Business mutation and outbox insert commit together. The outbox payload should contain the event id, aggregate id, type, and schema version.

Relay claim

Workers claim rows with a lease or `SKIP LOCKED`, publish, then mark success. Lease expiry makes crashes recoverable but permits duplicates.

Ordering

Partition by aggregate or sequence events per aggregate. Global ordering is expensive and usually unnecessary.

Retention

Delete or archive only after the retry and replay window. Keep enough metadata to diagnose and rebuild downstream projections.

Detailed boundary

co-committing intent with domain state

Operational consequence

relay retries and deduplication

Implementing it

Use a unique event id and idempotent consumers. Treat the relay as at-least-once.

Alert on oldest pending age, publish failures, lease expiry, and table growth.

Version event schemas and make replay a supported operation, not an emergency SQL script.

Exercise relay crash windows: acknowledge the broker, terminate the worker before the database mark, and verify the duplicate is harmless. Measure oldest unpublished age separately from broker lag.

Document retention, replay scope, and poison-message handling so an operator can recover a stuck event without editing business rows.

Use a monotonic aggregate sequence when consumers need per-entity order, and keep the event payload self-contained enough for replay after the source schema changes.

BEGIN;
UPDATE orders SET status = 'PAID' WHERE id = $1;
INSERT INTO outbox(id, type, aggregate_id, payload)
VALUES ($2, 'OrderPaid', $1, $3);
COMMIT;

Interview questions and how to answer them

What failure does the outbox close?

The crash window between a committed database change and its message publish, or between publish and recording that it was sent.

Can the relay publish duplicates?

Yes. A crash after broker acceptance and before marking the row sent causes a retry. Consumers must deduplicate.

How do you preserve order?

Use an aggregate sequence and partition or claim events in that order; do not assume concurrent relay workers preserve table order.

Why not use 2PC?

A broker may not participate, and 2PC can block. The outbox trades global atomicity for durable intent and idempotent eventual delivery.

What evidence would you inspect for transactional outbox?

Measure the boundary named in the design, compare it with the caller deadline and resource budget, and reproduce the contention or failure with more than one concurrent worker.

What is the tempting fix for this problem?

Changing a timeout, pool, or retry count alone usually moves the queue. First establish the invariant, then make the bounded mechanism and its failure outcome explicit.

Answers that lose the round

  • Publishing directly after a database commit
  • Marking dispatched before broker acknowledgement
  • Assuming one relay attempt means one delivery
  • Using a random event id on every retry
  • Deleting rows before downstream replay is possible
  • Putting remote calls inside the database transaction
  • Treating the local mechanism as a complete production guarantee
  • Changing the limit without measuring the resource it protects

Practise transactional outbox in a real repository

This repository tests the crash windows around a database write and event publish, including duplicate delivery and retrying a failed relay.

FAQ

Does the outbox remove eventual consistency?

No. It makes the delay and retry durable; consumers still apply the event asynchronously.

Where should deduplication live?

With the consumer’s side effect, ideally in the same transaction as applying that effect.

Can an outbox grow forever?

No. Retention, archiving, and replay policy must be explicit and monitored.

Related

More backend concepts