3.2 KiB
← Poison pills · Module README · Retry topics →
4. The dead-letter topic
The suffix is -dlt, not .DLT
public static final String RetryTopicConstants.DEFAULT_RETRY_SUFFIX = "-retry";
public static final String RetryTopicConstants.DEFAULT_DLT_SUFFIX = "-dlt";
Older Spring Kafka used .DLT, and most of the material online still says so. Getting it wrong
is not an exception — it is this, at WARN, once per record:
o.s.k.l.DeadLetterPublishingRecoverer : Destination resolver returned non-existent partition
payments-dlt-0, KafkaProducer will determine partition to use for this topic
[Producer] ... {payments-dlt=UNKNOWN_TOPIC_OR_PARTITION}
and then, on a cluster with auto-topic-creation disabled, the record is gone. Your safety net
dropped it and logged a warning. This module's tests were written against payments.DLT first
and failed exactly this way.
Two things follow: pre-create your DLT topics as part of provisioning, and alert on that WARN.
Same partition by default
DeadLetterPublishingRecoverer publishes to the same partition number as the original. If
your DLT has fewer partitions than the source topic, records from the high-numbered partitions
have nowhere to go. Either give the DLT the same partition count, or set
recoverer.setPartitionResolver((record, ex) -> null); // let the producer choose
The headers, and the one that will mislead you
From docs/output/retry-and-dlt.txt:
kafka_dlt-exception-fqcn org.springframework.kafka.listener.ListenerExecutionFailedException
kafka_dlt-exception-cause-fqcn com.ankurm.kafkaerrors.Failures$TransientFailure
kafka_dlt-original-topic payments
kafka_dlt-original-consumer-group payments
kafka_dlt-exception-fqcn is always the wrapper for a listener failure. Build a DLT triage
dashboard grouped by that header and every failure in the estate lands in one bucket called
ListenerExecutionFailedException. The field you want is kafka_dlt-exception-cause-fqcn.
(For a deserialization failure there is no wrapper, so the two headers agree. That inconsistency is worth knowing if you are writing a tool over them.)
kafka_dlt-original-consumer-group is the one that saves you when several groups consume the
same topic and share a DLT.
Replay
A DLT is only useful if you can put records back. The mechanics are a copy:
- read from
<topic>-dltwith a byte-array deserializer — the payload may be the thing that could not be deserialized - read
kafka_dlt-original-topicandkafka_dlt-original-consumer-groupto decide where it belongs and whether it is yours - republish to the original topic, stripping the
kafka_dlt-*headers so a second failure is not confused with the first - do it deliberately, in bounded batches, after the cause is fixed
Automatic replay is almost always wrong: the records are on the DLT precisely because something was not transient, and a loop that moves them back on a timer is a slow-motion outage. A replay you run by hand, having read the failure, is the tool worth building.