Distributed systems

Dead-letter queue filling silently

Written and reviewed by Sahil Srivastav

Distributed systemsQueuesOperations
consumer processed=100000 failed=0 dlq_depth=18432

What this error actually means

A dead-letter queue is a durable record of messages the normal consumer could not process. It is not a success sink. If its depth grows without an alert, the system is accepting loss of business work while dashboards report healthy throughput.

A poison message may fail deterministically, while transient failures may be exhausted by an overly short retry policy. Moving it to the DLQ changes ordering and often removes it from normal lag, so the source can look caught up while obligations accumulate elsewhere.

Recovery requires the original payload, headers, failure reason, and source position. A DLQ without those fields is an archive, not a repair mechanism.

Causes, most common first

  1. 1No alert on DLQ depth or age. The queue is treated as an implementation detail.
  2. 2Retry policy exhausts transient errors. Temporary dependency failures become permanent DLQ entries.
  3. 3Poison payload or schema mismatch. One message fails every attempt.

When you see it

  • Main queue lag is zero while DLQ depth grows
  • Failures are counted only in consumer logs
  • Operators replay messages manually from a partial export
  • A schema deploy causes a sudden DLQ slope

How to diagnose it

Step 1

Measure depth and oldest age

Count is insufficient; age tells you customer impact.

kafka-consumer-groups.sh --bootstrap-server broker:9092 --describe --group dlq-replayer
redis-cli ZCARD dead-letter

Step 2

Group failures by reason and version

Inspect headers, schema id, producer version, and exception class.

Step 3

Sample without mutating

Read a bounded sample and preserve raw payloads before replaying or deleting.

The fix

Alert on depth, oldest age, and growth rate.

Classify transient versus permanent failures and use a bounded retry schedule.

Retain original key, partition, offset, headers, payload, and stack trace.

Fix the consumer before replaying; replay into a quarantine topic if needed.

Deduplicate replay by event identity and preserve per-key ordering where required.

dlq.publish({ payload, key, sourceOffset, headers, error: str(exc), failedAt: now })

How to stop it coming back

  • Exercise poison messages in staging
  • Set retention based on recovery time
  • Provide a controlled replay tool
  • Dashboard DLQ age by tenant and reason
  • Require an owner for every DLQ

Practise production debugging in a real repository

Reading about a failure and reproducing one are different skills. Gronex ships broken backend repositories with failing test suites that encode the real invariant, so you debug from evidence instead of memorising symptoms.

FAQ

Can I delete old DLQ messages?

Only after a documented decision that the business work is obsolete or recovered elsewhere. Deletion is data loss.

Should DLQ messages be retried forever?

No. Separate transient retries from poison handling and alert humans when automated attempts are exhausted.

Why is zero main lag misleading?

Because the failed messages have been removed from the normal consumer path; DLQ depth is the missing workload.

Related

Other errors engineers hit next to this one

Full error and symptom index →