1
0
Files
spring-kafka-exactly-once-b…/docs/07-corner-cases.md
Ankur e23ca359eb Exactly-once with Spring Kafka on Spring Boot 4: demonstrations, docs and captured output
Companion suite for the ankurm.com guide. Every test runs against an in-JVM KRaft broker
(spring-kafka-test), so there is no Docker requirement. Verified on Spring Boot 4.1.0,
Spring Kafka 4.1.0, kafka-clients 4.2.1, JDK 25.0.3. results/full-run.txt is unedited
output of mvn test: 15 tests, 0 failures.

Demonstrations
  _01_idempotence   verifies idempotence from the BROKER via Admin.describeProducers rather
                    than by reading config back: a default producer registers one producer
                    state, an acks=1 producer registers zero. Also covers the ConfigException
                    you get only when idempotence is requested explicitly.
  _02_transactions  transaction markers consuming offsets, aborted records surviving in the
                    log, read_committed vs read_uncommitted, and a correction: markers do NOT
                    inflate consumer lag (a drained consumer reaches zero), they break
                    counting records by offset arithmetic (measured 100% overstatement).
  _03_read_process_write
                    the Spring loop with container-managed transactions; a listener that
                    fails after sending runs twice and is committed once.
  _04_fencing       shared transactional.id fences the zombie; a random per-startup id lets
                    both producers commit and silently defeats the guarantee entirely.
  _05_database      Kafka EOS does not extend to a database. Naive listener writes 2 rows for
                    1 message; the same listener keyed on topic-partition-offset writes 1.
  _06_performance   cost per transaction size, median of 5 interleaved rounds.

Docs
  Seven chapters: Boot 4 setup (auto-configuration moved to spring-boot-starter-kafka),
  idempotence, transactions and markers, fencing and KIP-890, the database boundary and the
  outbox pattern, performance, and corner cases.

Measured highlights
  acks=1 producer          0 broker-tracked producer states (idempotence silently off)
  5 records in one tx      endOffset 6; offset span overstates record count by 100%
  1 record  per tx            711 records/s
  100 records per tx       78,770 records/s  (~111x)
2026-07-31 22:27:35 +05:30

110 lines
5.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 7. Corner cases and things that bite
The blog post covers the main concepts. This chapter is the long tail — behaviours that do not fit a
narrative but do show up in production.
## 7.1 `transaction.timeout.ms` vs `max.poll.interval.ms`
The most common "it worked in dev" failure. Your listener must finish a whole poll batch, commit,
and get back to `poll()` inside **both** limits:
| Setting | Default | Owned by |
|---|---|---|
| `transaction.timeout.ms` | 60,000 (Spring sets it; Kafka's producer default is 60s) | producer |
| `transaction.max.timeout.ms` | 900,000 | **broker** — caps the above |
| `max.poll.interval.ms` | 300,000 | consumer |
If processing exceeds `transaction.timeout.ms`, the coordinator aborts your transaction and your
commit fails with `ProducerFencedException` / `InvalidProducerEpochException`. If it exceeds
`max.poll.interval.ms`, the consumer is ejected from the group and the partitions rebalance —
usually producing both symptoms at once and a very confusing log.
Rule: `transaction.timeout.ms` should comfortably exceed the worst-case time to process one batch
of `max.poll.records`, and the broker's `transaction.max.timeout.ms` must be at least as large or
producer creation fails outright.
## 7.2 A long transaction stalls other people's consumers
Covered in [03](03-transactions-and-markers.md) §3.5 and worth repeating because it is
counter-intuitive: a `read_committed` consumer cannot read past the Last Stable Offset, which is
held back by the *oldest open transaction on the partition* — regardless of which producer opened
it. One slow producer degrades every consumer on that partition.
## 7.3 `@KafkaListener` with manual acknowledgment does nothing useful
With a `KafkaTransactionManager` present, `AckMode.MANUAL` / `MANUAL_IMMEDIATE` and calls to
`Acknowledgment.acknowledge()` are irrelevant: the offset commit happens inside the transaction, at
the end of the batch. Leaving manual ack code in place is harmless but misleading — future readers
will believe it controls something.
## 7.4 `executeInTransaction` is not the same as being in a listener transaction
```java
kafkaTemplate.executeInTransaction(t -> t.send(TOPIC, value));
```
This starts a **local** producer transaction. It is the right way to produce transactionally from
outside a listener (a REST controller, a scheduled job, a test). Inside a listener that already has
a container-managed transaction, do **not** call it — plain `send()` enlists in the existing
transaction automatically. Nesting produces `IllegalStateException` or, worse, a second independent
transaction that commits separately.
## 7.5 Error handling interacts with transactions
`DefaultErrorHandler` retries in-process by default and then, if configured, publishes to a dead
letter topic via `DeadLetterPublishingRecoverer`. Two things to know:
- The DLT publish happens **inside the same transaction** when a `KafkaTransactionManager` is
present, so "record moved to DLT and offset advanced" is atomic. That is the behaviour you want,
and it is why a DLT is the correct terminal state for a poison record under EOS.
- Retries consume `max.poll.interval.ms`. A back-off of 10 attempts × 30 s will eject the consumer
from the group long before it gives up. Use `DefaultErrorHandler` with a bounded back-off and let
the DLT absorb the rest.
## 7.6 Producers are pooled, and `transactional.id` suffixes are allocated from that pool
Spring caches producers. With transactions, each concurrent transaction needs its own
`transactional.id`, so Spring allocates `prefix + suffix`. `TransactionIdSuffixStrategy`
(Spring Kafka 3.2+) controls that allocation; the default reuses suffixes as producers return to the
cache. If you see unexpected fencing under high concurrency, look here before looking at your code.
## 7.7 Consumer group `isolation.level` cannot be changed per-record
It is a consumer-level setting. If some records on a topic must be visible before commit and others
after, you need two topics, not two isolation levels.
## 7.8 Compacted topics and transactions
Transaction markers interact with log compaction: a marker cannot be cleaned until the transaction
it belongs to is complete, and a **hanging** transaction therefore blocks compaction of the whole
partition indefinitely. This was one of the motivations for
[KIP-890](https://cwiki.apache.org/confluence/display/KAFKA/KIP-890). If a compacted topic stops
shrinking, check for open transactions before checking your compaction settings.
## 7.9 Exactly-once across two Kafka clusters is not a thing
A transaction is scoped to one cluster's transaction coordinator. MirrorMaker 2 and most replication
tooling are at-least-once across clusters. A cross-cluster pipeline is at-least-once end to end no
matter what each side does internally.
## 7.10 Kafka Streams does this for you
If your topology is genuinely read-process-write over Kafka only, `processing.guarantee=exactly_once_v2`
in Kafka Streams gives you all of the above with none of the wiring, including state store changelogs
in the same transaction. Reach for Spring Kafka transactions when you need to interleave non-Kafka
work or you are not building a streaming topology — not because Streams cannot do it.
## 7.11 Things that are true and often assumed not to be
- **Idempotence and transactions do not require a special broker configuration** beyond replication
factors adequate for `__transaction_state`.
- **`enable.idempotence` is on by default** (Kafka 3.0+) — see [02](02-idempotence.md).
- **A transactional producer can send to topics it never mentioned up front.** Partitions are added
to the transaction dynamically as you send.
- **Reading your own writes** inside a transaction is not possible — a `read_committed` consumer
will not see them until commit, including in the same process.
---
Back to the [README](../README.md).