3.0 KiB
← Two kinds of failure · Module README · Poison pills →
2. What the default actually does
If you configure nothing, the container factory installs a DefaultErrorHandler with
SeekUtils.DEFAULT_BACK_OFF and a recoverer that logs. Run the back-off and read it off
(docs/output/default-backoff.txt):
=== DefaultErrorHandler default back-off ===
interval 0 ms
max attempts 9 retries
SeekUtils.DEFAULT_MAX_FAILURES = 10
retry intervals [0, 0, 0, 0, 0, 0, 0, 0, 0]
total deliveries 10
Ten deliveries, zero milliseconds apart, and then the record is dropped.
Both halves of that surprise people. It is not "retry with backoff" — it is ten immediate attempts as fast as the consumer thread can run them, which against a downstream that is overloaded is ten times the load at the worst moment. And "then dropped" means exactly that: the default recoverer logs the failure and the offset moves on. There is no dead-letter topic unless you make one.
Giving it a back-off and a destination
@Bean
DefaultErrorHandler errorHandler(KafkaOperations<String, Object> template) {
DeadLetterPublishingRecoverer recoverer = new DeadLetterPublishingRecoverer(template);
DefaultErrorHandler handler = new DefaultErrorHandler(recoverer, new FixedBackOff(1000L, 2));
handler.addNotRetryableExceptions(PermanentFailure.class);
return handler;
}
FixedBackOff(1000L, 2) is one delivery plus two retries. Measured
(docs/output/retry-and-dlt.txt):
=== transient failure ===
deliveries 3
gap between 1&2 1007 ms (FixedBackOff interval 1000)
and the classified permanent failure gets exactly one delivery before going to the DLT.
ExponentialBackOffWithMaxRetries is usually the better choice than FixedBackOff for a
transient downstream, because a fixed interval synchronises every consumer in the group into
retrying at the same instant.
The cost of blocking retries
DefaultErrorHandler retries on the consumer thread. For the whole back-off, that partition
processes nothing else. FixedBackOff(1000L, 2) is three seconds of a stalled partition per
failing record — fine. A one-minute exponential back-off over five attempts is five minutes, and
if failures are correlated you have a stalled consumer group, not a retry policy.
Two consequences worth planning for:
max.poll.interval.msis your ceiling. Default five minutes. Block longer than that between polls and the broker evicts the consumer from the group, triggering a rebalance — which usually makes things worse. A back-off schedule that can exceed it is a bug.- Ordering is preserved, which is the one thing blocking retries give you that retry topics do not.
That trade — ordering versus throughput under failure — is the real decision, and it is covered in chapter 5.