Skip to main content

Transactional Outbox with the Spring Modulith Event Publication Registry

A real dual-write failure, a real embedded Kafka broker, and the Spring Modulith event publication registry acting as a genuine transactional outbox — including replaying an incomplete publication for real.

Here is a method that looks completely reasonable:

var invoice = invoices.save(new Invoice(sku, amountCents, "ISSUED"));
kafka.send("invoices", invoice.getId().toString(), toJson(invoice));
return invoice;

Save to the database. Tell Kafka about it. Two lines, in order, and nothing about either line looks dangerous on its own. The danger is between them: these are two separate systems, with two separate commits, and nothing makes both happen or neither. If the process dies after the first line and before the second, or if Kafka is simply unreachable at that moment, the database says the invoice exists and nobody downstream ever finds out. This is the dual-write problem, and this post spends its first section reproducing it for real before fixing it — with a mechanism this series has already introduced for an unrelated reason, repurposed here as what it actually also is: a transactional outbox.

Versions used in this post. Spring Boot 4.1.1 · Spring Modulith 2.1.1 · spring-modulith-events-kafka 2.1.1 (same GA line as the rest of this series; its own <release> tag points at the 2.2.0-M2 milestone, verified against Maven Central’s maven-metadata.xml rather than trusted) · spring-kafka 4.1.1 / kafka-clients 4.2.1, both managed by the Spring Boot BOM · JDK 25 LTS. The Kafka broker in every captured transcript is a real, embedded one (@EmbeddedKafka) — no Docker, no mocks.

Two systems, one write, no coordinator

The companion repo’s NaiveBillingService is that exact method, kept around deliberately as the “before” picture — it is never wired into the module’s real API, only called directly from a test built to break it:

@Transactional
public Invoice saveInvoiceOnly(String sku, int amountCents) {
    return invoices.save(new Invoice(sku, amountCents, "ISSUED"));
}

public Invoice issueInvoiceTheNaiveWay(String sku, int amountCents) {
    var invoice = saveInvoiceOnly(sku, amountCents);
    kafka.send("invoices-naive", invoice.getId().toString(),
            invoice.getSku() + ":" + invoice.getAmountCents());
    return invoice;
}

NaiveBillingService.java

The two steps are separated on purpose: saveInvoiceOnly() has its own @Transactional boundary and returns — commits, fully — before issueInvoiceTheNaiveWay() ever reaches the kafka.send(...) line. Point spring.kafka.bootstrap-servers at a port with nothing listening and run it for real:

2026-10-04T00:52:56.429+05:30 ERROR 5271 --- [outbox] [           main] o.s.k.support.LoggingProducerListener    : Exception thrown when sending a message with key='1' and payload='WIDGET-1:1999' to topic invoices-naive:

org.apache.kafka.common.errors.TimeoutException: Topic invoices-naive not present in metadata after 2000 ms.

Exception thrown synchronously from issueInvoiceTheNaiveWay(): org.springframework.kafka.KafkaException: Send failed
Invoice count regardless: 1 (was 0). The database write already committed inside saveInvoiceOnly() before this exception was ever thrown -- the exception came from the Kafka call, several lines later, on an already-committed row.

Full transcript: output/00-dual-write-problem.txt

The invoice is in the database. Kafka never received anything. That’s the dual-write problem, and reproducing it needed nothing contrived — just two real systems, called one after the other, with no third thing making them agree.

A correction to the usual “fire-and-forget send() fails silently” story. KafkaTemplate.send() returns a CompletableFuture, and the folklore is that a caller who doesn’t block on it never sees a failure. Running it for real shows that’s only sometimes true: when the broker can’t be reached at all to fetch topic metadata, Spring’s KafkaTemplate throws synchronously, wrapping the Kafka client’s TimeoutException in a KafkaException — visible in the stack trace above, not asserted from memory. The narrower, genuinely silent case — broker reachable, topic known, but the specific produce request times out waiting for an acknowledgment afterward — only surfaces through the future or a callback, and needs a broker that accepts the connection and then stops responding, not one that was never there. This post’s transcript proves the first case; it does not reproduce the second, and says so rather than blurring the two together.
Two commits, one timeline, zero coordination app save(Invoice) — commits kafka.send(…) broker unreachable — send fails here Invoice row: committed, permanent, already returned to the caller. Nothing links these two boxes. A failure in the second box cannot undo the first, and nothing retries the second box unless the caller wrote that logic by hand.

What actually makes something an outbox

The fix is not “add a retry around the Kafka call”. A retry still runs after the database transaction has already committed, in a second, uncoordinated step — it narrows the failure window, it doesn’t close it. The transactional outbox pattern closes it differently: write the intent to notify someone into the same transaction as the business data, in the same database, and let a separate process deliver that intent afterward, retrying for as long as it takes.

This series already has that mechanism, introduced in the first post for a different reason: Spring Modulith’s event publication registry. Every @ApplicationModuleListener gets one row, written the moment ApplicationEventPublisher.publishEvent(...) is called — inside the same transaction as whatever business write triggered it, before any listener has run. That row is the outbox entry.

@Transactional
public Invoice issueInvoice(String sku, int amountCents) {
    var invoice = invoices.save(new Invoice(sku, amountCents, "ISSUED"));
    events.publishEvent(new InvoiceIssued(invoice.getId(), invoice.getSku(), invoice.getAmountCents()));
    return invoice;
}

BillingManagement.java

One method, one transaction, two things committed atomically: the Invoice row, and a row in the registry recording that InvoiceIssued needs to reach whoever is listening for it. Whatever happens next — a successful delivery, a crash one instruction later, a broker that’s down for ten minutes — cannot erase the fact that the intent to notify already made it to disk, atomically, with the invoice itself.

One local transaction, then a separate, retryable delivery one local DB transaction INSERT invoices (…) INSERT event_publication (…) commits together, or neither does afterward, async externalize to Kafka or run any listener retry if it fails The registry row is the outbox entry. A crash or a down broker after this point delays delivery — it cannot make the invoice and the intent to notify disagree with each other, because they were never two writes.

Externalizing to a real broker

Routing a published event to Kafka specifically — rather than just to another in-process listener — is one annotation, once spring-modulith-events-kafka is on the classpath:

@Externalized("invoices::#{#this.invoiceId}")
public record InvoiceIssued(Long invoiceId, String sku, int amountCents) {
}

InvoiceIssued.java

The string splits on ::: invoices is the Kafka topic, and #{#this.invoiceId} is a SpEL expression — evaluated against the event itself as #this — used as the message key. Verified against the compiled annotation (org.springframework.modulith.events.Externalized, 2.1.1): it is @Target(TYPE) with one aliased string attribute; the routing syntax itself is parsed at runtime by the externalization machinery, not encoded in the annotation’s shape.

Run it against a real embedded Kafka broker (no Docker, no mock producer):

Kafka record received -- topic=invoices key=1 value={"invoiceId":1,"sku":"GADGET-7","amountCents":4500}
EVENT_PUBLICATION_ARCHIVE rows for InvoiceIssued: 2
EVENT_PUBLICATION rows still outstanding for InvoiceIssued: 0

Full transcript: output/01-registry-safety-net.txt

The message really arrived, with the invoice’s own id as the routing key, exactly as the SpEL expression specifies. Two archive rows rather than one because this module’s tests wire a second listener on the same event (next section); with completion-mode=ARCHIVE turned on, both listeners’ rows moved out of the live EVENT_PUBLICATION table once they completed — the previous post in this series left that setting commented out specifically so its own transcripts showed the bare defaults; this module turns it on for real.

Going deeper: why this needs spring-modulith-events-kafka at all

A plain @ApplicationModuleListener that calls kafkaTemplate.send(...) itself would also work, and would still get outbox safety from the registry row written for that listener. What @Externalized adds is routing metadata as data — topic and key live on the event type, not buried in a listener method body — and the library handles the translation to Kafka’s producer API, including serialization, consistently across every externalized event in the application, rather than once per hand-written listener.

When the downstream fails: incomplete publications and replay

The outbox row solves “did we remember to try” forever, atomically. It does not make the eventual delivery succeed on the first attempt — a broker can still be briefly unreachable, a listener can still throw. What it buys is a durable, queryable record of exactly which notifications are still owed, and a real API for redelivering them: IncompleteEventPublications.resubmitIncompletePublicationsOlderThan(Duration) — verified against the compiled events-api 2.1.1 jar, not assumed from documentation prose.

To prove the retry path without needing to actually take a running broker down mid-test, the companion repo wires a second, deliberately flaky listener on the same event — test-only code, never part of the packaged application, standing in for any downstream dependency that can be briefly unavailable:

@ApplicationModuleListener
public void on(InvoiceIssued event) {
    int attempt = attempts.incrementAndGet();
    if (failNextAttempt) {
        failNextAttempt = false;
        throw new IllegalStateException(
                "simulated downstream failure on attempt " + attempt + " for invoice " + event.invoiceId());
    }
}

FlakyAuditListener.java (test source, not packaged)

Incomplete, not lost listener throws EVENT_PUBLICATION row COMPLETION_DATE = null COMPLETED resubmitIncompletePublicationsOlderThan(…) listener re-invoked, succeeds The row in the middle box is exactly what “incomplete” means: no completion date, whether the listener never ran yet or ran and failed. Nothing about the invoice itself is touched by any of this.
Incomplete FlakyAuditListener publications after the simulated failure: 1
Incomplete FlakyAuditListener publications after resubmission: 0
Invoice 1 never changed -- the retry was purely about redelivering the event, the invoice row was correct the whole time.

Full transcript: output/02-replay-incomplete-publications.txt

One call to resubmitIncompletePublicationsOlderThan(Duration.ZERO) re-invoked the listener for every row still missing a completion date. This time it succeeded, and the row was marked complete. In a real deployment this call is what a scheduled job runs periodically — nobody has to notice the failure and react by hand, because the registry already knows exactly which rows are still owed.

Resubmission re-runs the listener, in full, from the start. If a listener does two things and the first one isn’t idempotent — charges a card, decrements stock — a resubmission after a partial failure can repeat that side effect. The registry guarantees the notification is attempted until it succeeds; it says nothing about what the listener does being safe to attempt twice. Making a listener idempotent is a separate, deliberate design decision, and the subject of the Saga post later in this series.
Going deeper: FailedEventPublications, and the companion ResubmissionOptions

IncompleteEventPublications covers any row with no completion date, whether or not it ever ran. A sibling interface, FailedEventPublications (since Spring Modulith 2.0), narrows that to rows a listener actually threw on, via resubmit(ResubmissionOptions). ResubmissionOptions.defaults() can be customized with withBatchSize(...), withMinAge(...), and withFilter(...) — all verified against the compiled class, all real methods on a real class, useful once a production system has enough outstanding publications that resubmitting all of them at once isn’t the right default.

What this still doesn’t give you

An outbox fixes the dual-write problem. It does not fix everything adjacent to it:

  • At-least-once, not exactly-once. A crash between a listener completing its work and the registry recording that completion means the resubmission job will run that listener again. Nothing in this pattern makes a second delivery impossible — only idempotent handling on the receiving end does.
  • Ordering across listeners is not guaranteed. The companion post’s event-cascade section already showed independent @ApplicationModuleListener methods running on separate threads with no ordering promise between them; externalizing one of them to Kafka doesn’t add an ordering guarantee that wasn’t there before.
  • The registry table is now part of your data model. It needs the same backup, monitoring, and growth planning as any other table that writes on every business transaction — completion-mode=ARCHIVE here is that planning, not a default you get for free.

None of these are reasons not to use an outbox. They’re the reasons “we added an outbox” is not the same claim as “our events can never be duplicated or lost in any order” — a distinction worth being precise about before it ships.

Should you reach for this? If a business transaction needs to reliably notify something outside your database — another service, a message broker, a webhook — and “reliably” has to survive a crash at the worst possible moment, yes, and Spring Modulith’s event publication registry gives you most of it for the cost of one annotation and a completion-mode setting. If nothing downstream cares whether a notification arrives, or a short delay is genuinely fine either way, this is machinery you don’t need yet.

Further reading

No Comments yet!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.