Skip to main content

Saga Pattern in Spring Boot: Orchestration vs Choreography with Kafka

The same order/payment/inventory saga built twice against a real Kafka broker — choreography and orchestration — with a real compensating transaction and a real duplicate-delivery idempotent-consumer test.

An order needs three things to happen: a payment reserved, inventory reserved, and the order itself confirmed. All three live in different services. A single database transaction can’t reach across all three — there’s no one database to commit against — so the question this post answers is: when the second step fails, who undoes the first one, and how do they know to?

That’s the whole problem a saga solves. Not a performance technique, not a message format — a way of getting a business operation that spans several services to end up consistent, using nothing but local transactions and a sequence of compensating steps when one of them fails.

Versions used in this post. Spring Boot 4.1.1 (verified GA against Maven Central’s maven-metadata.xml — 4.2.0-M2 is the newest entry there but is a milestone, not a release) · Spring Kafka 4.1.1, managed by the Boot BOM · JDK 25 LTS. Every transcript in this post comes from a real embedded Kafka broker (@EmbeddedKafka) — no Docker, no mocks, and no actual network calls between separate services: both coordination styles run as plain Spring beans inside one JVM, talking only over real Kafka topics, which is enough to prove how the coordination itself behaves without standing up three deployable services.

What a saga actually is, and why “just use a transaction” doesn’t work here

A relational database gives you atomicity for free: a transaction either fully commits or fully rolls back, and nothing in between is ever visible. The moment a business operation touches a second database or a second service, that guarantee is gone — there’s no shared log for two independent systems to agree on, which is exactly the dual-write problem this series already reproduced for real with a single service and its own message broker. A saga is what you reach for once the problem is no longer “my service and a broker” but “three separate services, each with the final say over its own data”.

Chris Richardson’s definition, the one most of the Java ecosystem cites, is blunt about the trade: “If a local transaction fails because it violates a business rule then the saga executes a series of compensating transactions that undo the changes that were made by the preceding local transactions.” There is no automatic rollback waiting in the wings — “a developer must design compensating transactions that explicitly undo changes made earlier in a saga rather than relying on the automatic rollback feature of ACID transactions.” Every line of Java in this post’s companion repo exists because of that one sentence.

One business operation, three separate local transactions Payment service reserve payment Inventory service reserve stock Order service confirm order Each box commits on its own, to its own database. Nothing wraps all three in one transaction — if the third box can’t happen, undoing the first two is a saga’s job, not a rollback’s.

Two structurally different ways exist to run that sequence and to trigger the compensations when it breaks: choreography, where “each local transaction publishes domain events that trigger local transactions in other services”, and orchestration, where “an orchestrator (object) tells the participants what local transactions to execute.” This post builds the same saga both ways, against the same real broker, and runs the same failure through each so the difference is something you can watch happen rather than take on faith.

Choreography: every participant decides for itself

The companion repo’s choreography saga has four small classes, and no class among them knows the saga exists as a whole. ChoreographySagaStarter creates the order and publishes one event; everything after that is a listener reacting to the event in front of it:

@KafkaListener(topics = Topics.CHOREO_ORDER_CREATED, groupId = "choreo-payment")
public void onOrderCreated(OrderCreated event) {
    payments.reserve(event.orderId(), event.amountCents());
    kafka.send(Topics.CHOREO_PAYMENT_RESERVED,
            new PaymentReserved(event.orderId(), event.sku(), event.qty()));
}

PaymentChoreographyListener.java

InventoryChoreographyListener reacts to that event the same way, and is the only class in the choreography saga with an opinion about whether stock is sufficient. Run it with enough stock on hand and the whole chain completes with nobody at the center:

Order 1 status: CONFIRMED
Stock remaining for GADGET-1: 8
Payment record status: RESERVED

Full transcript: output/00-choreography-happy-path.txt

Choreography: a chain of listeners, no single class aware of the whole Payment listener Inventory listener Order listener PaymentReserved InventoryReserved Each box’s own @KafkaListener method is the only place that step’s decision is made. Remove the arrows and ask “what runs this saga end to end” — there’s no box that answers that question. That’s choreography’s whole trade: no coordinator to build, but no single place to read either.

Choreography’s intermediate cost shows up the moment a saga grows past three steps: tracing “why was this order cancelled” means reading every listener’s subscription list and mentally re-assembling the chain, because no file states the sequence anywhere. For a saga this size that’s a few minutes; for a real checkout flow with six or seven participants, it’s the reason teams reach for distributed tracing just to answer a question a single orchestrator class would answer by existing.

When something fails: the compensating transaction

Run the same chain with too little stock on hand, and the saga has to undo the payment it already reserved. Nothing new coordinates this — PaymentChoreographyListener was already subscribed to the rejection event for exactly this reason:

@KafkaListener(topics = Topics.CHOREO_INVENTORY_REJECTED, groupId = "choreo-payment-compensation")
public void onInventoryRejected(InventoryRejected event) {
    payments.refund(event.orderId());
    kafka.send(Topics.CHOREO_PAYMENT_REFUNDED, new PaymentRefunded(event.orderId()));
}

PaymentChoreographyListener.java

Order 1 status: CANCELLED
Stock remaining for GADGET-1 (untouched by the rejected reservation): 1
Payment record status: REFUNDED

Full transcript: output/01-choreography-compensation.txt

The same listener, reused for the undo Inventory listener (rejects) Payment listener (refunds) Order listener (cancels) InventoryRejected PaymentRefunded PaymentChoreographyListener has two @KafkaListener methods, and the compensation is simply the second one — it exists because this class is already subscribed, not because anything told it to run.
Compensation undoes the effect, not the history. The refund is a new, forward-moving operation (payments.refund()), not a reversal of the original INSERT — the PaymentRecord row still exists, now marked REFUNDED. Chris Richardson’s point about the “lack of automatic rollback” is exactly this: nothing erases the fact that a reservation briefly existed, and any compensating transaction you write has to be correct about that, not just about restoring the final numbers.

Orchestration: one class makes every decision

The orchestration version runs the identical two scenarios, through a structurally different set of classes. PaymentOrchestrationHandler and InventoryOrchestrationHandler never decide what happens after they act — they carry out one command each and reply. Every decision lives in a single class instead:

@KafkaListener(topics = Topics.ORCH_REPLY_INVENTORY, groupId = "orchestrator")
public void onInventoryReply(InventoryReply reply) {
    if (reply.success()) {
        orders.confirm(reply.orderId());
        return;
    }
    kafka.send(Topics.ORCH_CMD_REFUND_PAYMENT,
            new RefundPaymentCommand(UUID.randomUUID().toString(), reply.orderId()));
}

OrderSagaOrchestrator.java

[happy path] Order 1 status: CONFIRMED
[happy path] Stock remaining for GADGET-ORCH-OK: 8
[happy path] Payment record status: RESERVED
[compensation] Order 2 status: CANCELLED
[compensation] Stock remaining for GADGET-ORCH-FAIL (untouched): 1
[compensation] Payment record status: REFUNDED

Full transcript: output/02-orchestration-saga.txt

Orchestration: one hub, every decision in one place OrderSagaOrchestrator decides every next step Payment handler Inventory handler cmd.reserve-payment cmd.reserve-inventory Solid arrows are commands out, dashed arrows are replies back. Neither handler ever talks to the other.

Lay this diagram next to the choreography one above: same three participants, same two outcomes, and the only structural difference is where the decision “what happens after inventory answers” lives. Orchestration puts it in one file you can read top to bottom; choreography spreads it across every listener’s subscription list. Neither is a correctness difference — both versions’ transcripts show identical outcomes — it’s entirely about where you’ll look when something goes wrong at 2 a.m.

Why duplicates are dangerous, and the idempotent consumer pattern

Kafka’s delivery guarantee is at-least-once by default — this series’ own kafka-basics module established that the default acknowledgement mode is BATCH, which is exactly why a consumer rebalance, a retry, or a producer resend can redeliver a message that already did its job. For most of this saga that’s harmless: replaying InventoryReserved a second time wouldn’t do anything, because nothing listens for it twice in a way that matters. A command that decrements real stock is a different story — carrying it out twice is a real bug, not a no-op.

InventoryOrchestrationHandler is the one listener in this saga where that matters, so it checks before it does anything else:

@KafkaListener(topics = Topics.ORCH_CMD_RESERVE_INVENTORY, groupId = "orch-inventory")
public void onReserveInventory(ReserveInventoryCommand command) {
    if (!idempotencyGuard.claim(command.commandId())) {
        return;
    }
    boolean reserved = inventory.tryReserve(command.sku(), command.qty());
    String reason = reserved ? null : "insufficient stock for " + command.sku();
    kafka.send(Topics.ORCH_REPLY_INVENTORY,
            new InventoryReply(command.commandId(), command.orderId(), reserved, reason));
}

InventoryOrchestrationHandler.java

claim() is the whole pattern, and it’s deliberately not clever:

@Transactional(propagation = Propagation.REQUIRES_NEW)
public boolean claim(String commandId) {
    try {
        processed.save(new ProcessedCommand(commandId));
        return true;
    } catch (DataIntegrityViolationException alreadyClaimed) {
        return false;
    }
}

IdempotencyGuard.java

The UNIQUE constraint on commandId — not an if statement in application code — is what makes this safe if two threads ever race on the exact same command: at most one insert can win. Publish the same command, same id, twice in a row:

ProcessedCommand rows for commandId f07ec35c-853c-4405-a024-b62d304dfabb: 1
Stock remaining for GADGET-DUP after two deliveries of the same command: 7
InventoryReply messages actually published: 1
  reply: InventoryReply[commandId=f07ec35c-853c-4405-a024-b62d304dfabb, orderId=999, success=true, reason=null]

Full transcript: output/03-idempotent-consumer-duplicate.txt

Same commandId, delivered twice delivery #1 claim() succeeds stock reserved, reply sent delivery #2 (duplicate) claim() fails: UNIQUE returns — nothing else runs 10 – 3 = 7 7 (untouched)
This guards the command, not the whole saga. If a listener does two things — decrement stock, then also update a shipping estimate — a crash between them and a retry would still only re-run the whole method once more, because the guard is keyed on the command, not on each individual side effect inside it. A listener with multiple side effects still needs each of them to tolerate being retried together, or needs its own finer-grained idempotency keys.

Choreography vs orchestration: a failure matrix

Both versions of this saga produce identical outcomes for identical inputs — every transcript above proves that. What actually differs between them only shows up once something goes wrong, or once the saga has to grow:

SituationChoreographyOrchestration
A participant never replies (crashes mid-step)Every other participant is still waiting on an event that will never arrive; nothing times out unless each listener separately implements oneThe orchestrator is the one place a timeout belongs, and it already knows exactly which step is outstanding
Adding a fourth participant to the sagaNew listener, subscribed to an existing event — no existing class changesThe orchestrator’s code changes to know about the new step; every existing participant stays untouched
“Why was this order cancelled?”Reconstructed by reading every listener’s subscriptions in turn, or from a trace spanning several servicesAnswered by one class’s logic, in order, in one file
A duplicate/redelivered messageExactly as dangerous as in orchestration — the idempotent-consumer pattern above applies to either style equallySame
Single point of failureNone of these classes is more critical than another — but there’s also no one place to add saga-wide monitoringThe orchestrator becomes the thing that must stay up and be deployed carefully; it’s also the one dashboard you’d build first

Richardson’s own framing is the simplest way to decide between them: choreography costs you less up front and avoids a central dependency, right up until the number of participants makes “what happens next” something you can no longer hold in your head. Three participants, as in this post’s demo, is comfortably on the choreography side of that line for most teams; six or seven participants with multiple failure branches is usually where orchestration’s one readable file starts winning.

Should you reach for a saga at all? Only once a business operation genuinely spans more than one service’s own data, with no way to make it one local transaction. If everything the operation touches lives in a single service’s own database, that’s a local transaction, full stop — sagas trade ACID guarantees for reach, and you only need that trade when the reach is unavoidable. Spring Kafka’s own KafkaTransactionManager can’t make this problem go away either: chaining it with a database transaction manager commits the database transaction, then the Kafka transaction, as two separate commits in sequence — a narrower window than no coordination at all, but not the same guarantee a saga’s explicit compensations give you.

Further reading

No Comments yet!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.