Distributed systems
Dead-letter queue filling silently
Written and reviewed by Sahil Srivastav
consumer processed=100000 failed=0 dlq_depth=18432What this error actually means
A dead-letter queue is a durable record of messages the normal consumer could not process. It is not a success sink. If its depth grows without an alert, the system is accepting loss of business work while dashboards report healthy throughput.
A poison message may fail deterministically, while transient failures may be exhausted by an overly short retry policy. Moving it to the DLQ changes ordering and often removes it from normal lag, so the source can look caught up while obligations accumulate elsewhere.
Recovery requires the original payload, headers, failure reason, and source position. A DLQ without those fields is an archive, not a repair mechanism.
Causes, most common first
- 1No alert on DLQ depth or age. The queue is treated as an implementation detail.
- 2Retry policy exhausts transient errors. Temporary dependency failures become permanent DLQ entries.
- 3Poison payload or schema mismatch. One message fails every attempt.
When you see it
- Main queue lag is zero while DLQ depth grows
- Failures are counted only in consumer logs
- Operators replay messages manually from a partial export
- A schema deploy causes a sudden DLQ slope
How to diagnose it
Step 1
Measure depth and oldest age
Count is insufficient; age tells you customer impact.
kafka-consumer-groups.sh --bootstrap-server broker:9092 --describe --group dlq-replayer
redis-cli ZCARD dead-letterStep 2
Group failures by reason and version
Inspect headers, schema id, producer version, and exception class.
Step 3
Sample without mutating
Read a bounded sample and preserve raw payloads before replaying or deleting.
The fix
Alert on depth, oldest age, and growth rate.
Classify transient versus permanent failures and use a bounded retry schedule.
Retain original key, partition, offset, headers, payload, and stack trace.
Fix the consumer before replaying; replay into a quarantine topic if needed.
Deduplicate replay by event identity and preserve per-key ordering where required.
dlq.publish({ payload, key, sourceOffset, headers, error: str(exc), failedAt: now })How to stop it coming back
- Exercise poison messages in staging
- Set retention based on recovery time
- Provide a controlled replay tool
- Dashboard DLQ age by tenant and reason
- Require an owner for every DLQ
FAQ
Can I delete old DLQ messages?
Only after a documented decision that the business work is obsolete or recovered elsewhere. Deletion is data loss.
Should DLQ messages be retried forever?
No. Separate transient retries from poison handling and alert humans when automated attempts are exhausted.
Why is zero main lag misleading?
Because the failed messages have been removed from the normal consumer path; DLQ depth is the missing workload.
Related
Other errors engineers hit next to this one
- pg client already connected or released twice
- Sequelize / Knex pool acquire timeout
- ERR_STREAM_PREMATURE_CLOSE during an upload
- Process exits before asynchronous writes finish
- ERR_MODULE_NOT_FOUND during ESM/CommonJS migration
- Unbounded JSON body blocks the loop or exhausts memory
- Undici / fetch connections remain occupied
- ERR_UNHANDLED_REJECTION in a worker thread