Distributed systems

Dead letter queue: interview questions and how to answer them

A dead-letter queue parks messages a consumer cannot process so the rest of the stream keeps flowing — which only helps if something distinguishes permanent failures from transient ones and someone actually drains it.

Written and reviewed by Sahil Srivastav

Distributed systemsMessagingOperational design

What it actually is

A dead-letter queue is a side channel for messages that have exhausted their retries. Without one, a message the consumer cannot handle is retried forever: in a FIFO or partition-ordered stream it blocks everything behind it, and in an unordered queue it burns capacity indefinitely. The DLQ converts an unbounded failure into a bounded one by moving the message aside.

The classification it depends on is the substance of the design. A transient failure — a timeout, a 503, a deadlock, a momentarily unreachable dependency — should be retried, because the same message will succeed later. A permanent failure — a schema violation, a reference to an entity that will never exist, a business rule rejection — will fail identically on every attempt, so retrying it is pure waste and it should go to the DLQ on the first or second try.

What a DLQ is not is a bin for errors you have not thought about. Messages that land there are real work that did not happen: an order not fulfilled, a payment not reconciled, a notification not sent. The queue is a durable record of incomplete business, which is why an unmonitored DLQ is a liability rather than a safety net.

Why it matters in production

Because a single malformed message can halt a pipeline. With per-key ordering, one poison message blocks its partition or message group entirely; with SQS FIFO it blocks the group until the visibility timeout expires, repeatedly. The DLQ is what keeps one bad record from becoming an outage for every other record.

It also matters because retry storms are self-inflicted overload. A dependency that is failing for everyone receives the full retry volume of every consumer at once, and infinite retries on permanent errors add load that can never succeed. Capping attempts and diverting to a DLQ is part of protecting the dependency, not just the consumer.

And it is a favourite interview area because the follow-ups are operational: who is paged when the DLQ is non-empty, how a message is replayed, whether replay is safe, and what happens to ordering. Those questions separate people who have run a pipeline from people who have configured one.

How it works

Classify before you retry

The handler should decide, per exception type, whether the failure is retryable. A validation error or a 4xx from a downstream is permanent — fail fast to the DLQ. A timeout, a connection reset, a 429, a 5xx, or a serialisation conflict is transient — retry with backoff. Retrying everything uniformly wastes the retry budget on messages that cannot succeed and delays the ones that can.

Retry policy shape

Exponential backoff with jitter, a bounded attempt count, and a bounded total age. The age bound matters independently: a message retried for six hours may be worthless by the time it succeeds, and a downstream that comes back to a thundering herd of hours-old retries can fail again immediately. In-process retries handle short blips; requeue-with-delay handles longer ones without holding a consumer slot.

Preserve the diagnosis with the message

A DLQ entry with only the payload is nearly useless. Attach the failure reason, the exception class and message, the attempt count, the first and last failure timestamps, the consumer version, and the trace id. Investigation then starts from the message rather than from a log search across a retention window that may have expired.

Redrive is not free

Replaying a DLQ means delivering messages again, possibly after a code fix and possibly days later. Two consequences: the consumer must be idempotent, because some of those messages may have partially applied before failing; and the dedup or idempotency window must exceed the maximum DLQ age, or a replayed message is treated as new and reapplied. Redrive into a separate queue with a lower rate limit avoids overwhelming a freshly repaired dependency.

Ordering is sacrificed deliberately

Parking a message means the messages after it are processed first, so per-key order is broken for that key. If order is a hard invariant, the correct behaviour is to stall the key rather than skip it — park the whole key, not the message — which means the DLQ design and the ordering design have to be decided together rather than separately.

Implementing it

Alert on DLQ depth greater than zero and on age of oldest message, not on a threshold like a hundred. One dead-lettered payment is an incident; discovering it a week later because the alert fired at a hundred is much worse.

Make the DLQ per consumer, not shared. A shared DLQ mixes failures with different causes and owners, and makes selective redrive after a targeted fix impractical.

Build the redrive path before you need it, with a dry run, a filter, and a rate limit. Manually re-publishing from a console during an incident is how messages get double-applied.

Track the DLQ arrival rate as a product metric. A steady trickle usually means an unhandled legitimate case in the schema or the business rules — the DLQ is telling you about a requirements gap, not an infrastructure problem.

Interview questions and how to answer them

Why do you need a dead-letter queue at all?

To stop one unprocessable message from halting or degrading the whole stream. In an ordered stream it blocks its partition; in an unordered one it consumes capacity forever. Parking it bounds the damage and turns the failure into a visible, actionable record rather than an endless retry loop.

How do you decide whether to retry or dead-letter?

By classifying the failure. Transient — timeout, connection reset, 429, 5xx, deadlock, optimistic-lock conflict — gets retried with exponential backoff and jitter, because the same input will succeed later. Permanent — schema violation, unparseable payload, 4xx from a downstream, a business rule that will never pass — goes straight to the DLQ, because every retry produces the identical failure. Anything unclassified I would treat as transient but with a low attempt cap.

You have 12,000 messages in a DLQ after a bug fix. How do you replay them?

First confirm the consumer is idempotent, since some messages may have partially applied. Then redrive into a dedicated replay queue rather than the live one, rate-limited so the downstream is not flooded, starting with a small sample verified end to end. Check that the idempotency window covers the age of the oldest message; if it does not, dedup must come from a natural business key instead. Keep the original DLQ contents until the replay is verified.

What information do you put on a dead-lettered message?

Exception class and message, the failing stage or handler, attempt count, first-seen and last-attempt timestamps, the consumer build or version, the trace id, and the original message metadata including partition and offset. The goal is that an engineer can decide “code bug”, “bad data”, or “dependency was down” from the DLQ record alone.

Does a DLQ break your ordering guarantee?

Yes — parking a message lets later messages for the same key proceed, so the key’s sequence is violated. If ordering is a real invariant, the right response to a failure is to stall the key: stop consuming for that key and alert, rather than skipping ahead. That trades throughput for correctness, and it is a decision to make explicitly rather than discover.

What alerting would you put on a DLQ?

Depth greater than zero for high-value streams, age of the oldest message, and arrival rate. Age matters because a message sitting for days may fall outside dedup windows or business deadlines. Arrival rate matters because a steady trickle signals an unhandled legitimate case rather than a one-off, and that is a backlog item rather than a page.

Answers that lose the round

  • Retrying every failure the same way, so permanent errors consume the whole retry budget and delay recoverable work
  • Leaving retries unbounded, which lets one poison message block a partition or message group indefinitely
  • Writing to the DLQ without the failure reason, attempt count, or trace id, making investigation guesswork
  • Treating the DLQ as a dumping ground with no owner and no alert — dead-lettered messages are unfinished business
  • Redriving without idempotency, so partially applied messages are reapplied
  • Having a dedup window shorter than the maximum DLQ age, so replayed messages look new
  • Alerting only on a large depth threshold, which hides the single high-value message that failed

Practise in a real repository

Explaining a concept and enforcing it in code are different skills, and machine coding rounds test the second. Gronex ships broken backend repositories whose test suites assert the invariant rather than the happy path.

FAQ

Should a DLQ have a retention limit?

It needs one, but a long one — typically the maximum you can store, because deleting a dead-lettered message deletes unfinished business. The right control is not aggressive retention but an alert that forces the queue to be drained long before retention becomes relevant.

Does Kafka have dead-letter queues?

Not natively at the broker level; the pattern is implemented in the consumer or framework. Kafka Connect has `errors.deadletterqueue.topic.name`, and Spring Kafka or a hand-rolled handler publishes failures to a designated topic. Because Kafka consumers control their own offsets, the consumer must also commit past the failed record after publishing it, or it will simply re-consume the same message.

Is a DLQ the same as a retry queue?

No, and separating them is good practice. A retry queue holds messages that will be attempted again automatically, often with a delay; a DLQ holds messages that will not be retried without human action. Using one queue for both means either automatic retries of hopeless messages or manual handling of transient blips.

Related

More backend concepts