An order needs three things to happen: a payment reserved, inventory reserved, and the order itself confirmed. All three live in different services. A single database transaction can’t reach across all three — there’s no one database to commit against — so the question this post answers is: when the second step fails, who undoes the first one, and how do they know to?
That’s the whole problem a saga solves. Not a performance technique, not a message format — a way of getting a business operation that spans several services to end up consistent, using nothing but local transactions and a sequence of compensating steps when one of them fails.
Versions used in this post. Spring Boot 4.1.1 (verified GA against Maven Central’s maven-metadata.xml —4.2.0-M2is the newest entry there but is a milestone, not a release) · Spring Kafka 4.1.1, managed by the Boot BOM · JDK 25 LTS. Every transcript in this post comes from a real embedded Kafka broker (@EmbeddedKafka) — no Docker, no mocks, and no actual network calls between separate services: both coordination styles run as plain Spring beans inside one JVM, talking only over real Kafka topics, which is enough to prove how the coordination itself behaves without standing up three deployable services.
What a saga actually is, and why “just use a transaction” doesn’t work here
A relational database gives you atomicity for free: a transaction either fully commits or fully rolls back, and nothing in between is ever visible. The moment a business operation touches a second database or a second service, that guarantee is gone — there’s no shared log for two independent systems to agree on, which is exactly the dual-write problem this series already reproduced for real with a single service and its own message broker. A saga is what you reach for once the problem is no longer “my service and a broker” but “three separate services, each with the final say over its own data”.
Chris Richardson’s definition, the one most of the Java ecosystem cites, is blunt about the trade: “If a local transaction fails because it violates a business rule then the saga executes a series of compensating transactions that undo the changes that were made by the preceding local transactions.” There is no automatic rollback waiting in the wings — “a developer must design compensating transactions that explicitly undo changes made earlier in a saga rather than relying on the automatic rollback feature of ACID transactions.” Every line of Java in this post’s companion repo exists because of that one sentence.
Two structurally different ways exist to run that sequence and to trigger the compensations when it breaks: choreography, where “each local transaction publishes domain events that trigger local transactions in other services”, and orchestration, where “an orchestrator (object) tells the participants what local transactions to execute.” This post builds the same saga both ways, against the same real broker, and runs the same failure through each so the difference is something you can watch happen rather than take on faith.
- Saga pattern, primary source: microservices.io — Saga
- A worked example of both styles: microservices.io — A tour of two sagas
Choreography: every participant decides for itself
The companion repo’s choreography saga has four small classes, and no class among them knows
the saga exists as a whole. ChoreographySagaStarter creates the order and publishes
one event; everything after that is a listener reacting to the event in front of it:
@KafkaListener(topics = Topics.CHOREO_ORDER_CREATED, groupId = "choreo-payment")
public void onOrderCreated(OrderCreated event) {
payments.reserve(event.orderId(), event.amountCents());
kafka.send(Topics.CHOREO_PAYMENT_RESERVED,
new PaymentReserved(event.orderId(), event.sku(), event.qty()));
}
PaymentChoreographyListener.java
InventoryChoreographyListener reacts to that event the same way, and is the only
class in the choreography saga with an opinion about whether stock is sufficient. Run it with
enough stock on hand and the whole chain completes with nobody at the center:
Order 1 status: CONFIRMED
Stock remaining for GADGET-1: 8
Payment record status: RESERVED
Full transcript: output/00-choreography-happy-path.txt
Choreography’s intermediate cost shows up the moment a saga grows past three steps: tracing “why was this order cancelled” means reading every listener’s subscription list and mentally re-assembling the chain, because no file states the sequence anywhere. For a saga this size that’s a few minutes; for a real checkout flow with six or seven participants, it’s the reason teams reach for distributed tracing just to answer a question a single orchestrator class would answer by existing.
- Going deeper on the tracing cost specifically: microservices.io’s worked comparison of both styles
When something fails: the compensating transaction
Run the same chain with too little stock on hand, and the saga has to undo the payment it
already reserved. Nothing new coordinates this — PaymentChoreographyListener was
already subscribed to the rejection event for exactly this reason:
@KafkaListener(topics = Topics.CHOREO_INVENTORY_REJECTED, groupId = "choreo-payment-compensation")
public void onInventoryRejected(InventoryRejected event) {
payments.refund(event.orderId());
kafka.send(Topics.CHOREO_PAYMENT_REFUNDED, new PaymentRefunded(event.orderId()));
}
PaymentChoreographyListener.java
Order 1 status: CANCELLED
Stock remaining for GADGET-1 (untouched by the rejected reservation): 1
Payment record status: REFUNDED
Full transcript: output/01-choreography-compensation.txt
Compensation undoes the effect, not the history. The refund is a new, forward-moving operation (payments.refund()), not a reversal of the original INSERT — thePaymentRecordrow still exists, now markedREFUNDED. Chris Richardson’s point about the “lack of automatic rollback” is exactly this: nothing erases the fact that a reservation briefly existed, and any compensating transaction you write has to be correct about that, not just about restoring the final numbers.
Orchestration: one class makes every decision
The orchestration version runs the identical two scenarios, through a structurally different
set of classes. PaymentOrchestrationHandler and InventoryOrchestrationHandler
never decide what happens after they act — they carry out one command each and reply. Every
decision lives in a single class instead:
@KafkaListener(topics = Topics.ORCH_REPLY_INVENTORY, groupId = "orchestrator")
public void onInventoryReply(InventoryReply reply) {
if (reply.success()) {
orders.confirm(reply.orderId());
return;
}
kafka.send(Topics.ORCH_CMD_REFUND_PAYMENT,
new RefundPaymentCommand(UUID.randomUUID().toString(), reply.orderId()));
}
[happy path] Order 1 status: CONFIRMED
[happy path] Stock remaining for GADGET-ORCH-OK: 8
[happy path] Payment record status: RESERVED
[compensation] Order 2 status: CANCELLED
[compensation] Stock remaining for GADGET-ORCH-FAIL (untouched): 1
[compensation] Payment record status: REFUNDED
Full transcript: output/02-orchestration-saga.txt
Lay this diagram next to the choreography one above: same three participants, same two outcomes, and the only structural difference is where the decision “what happens after inventory answers” lives. Orchestration puts it in one file you can read top to bottom; choreography spreads it across every listener’s subscription list. Neither is a correctness difference — both versions’ transcripts show identical outcomes — it’s entirely about where you’ll look when something goes wrong at 2 a.m.
- A worked orchestration-based saga, step by step: microservices.io — Implementing an orchestration-based saga
Why duplicates are dangerous, and the idempotent consumer pattern
Kafka’s delivery guarantee is at-least-once by default — this series’ own
kafka-basics module established that the default acknowledgement mode is
BATCH, which is exactly why a consumer rebalance, a retry, or a producer resend can
redeliver a message that already did its job. For most of this saga that’s harmless: replaying
InventoryReserved a second time wouldn’t do anything, because nothing listens for
it twice in a way that matters. A command that decrements real stock is a different
story — carrying it out twice is a real bug, not a no-op.
InventoryOrchestrationHandler is the one listener in this saga where that matters,
so it checks before it does anything else:
@KafkaListener(topics = Topics.ORCH_CMD_RESERVE_INVENTORY, groupId = "orch-inventory")
public void onReserveInventory(ReserveInventoryCommand command) {
if (!idempotencyGuard.claim(command.commandId())) {
return;
}
boolean reserved = inventory.tryReserve(command.sku(), command.qty());
String reason = reserved ? null : "insufficient stock for " + command.sku();
kafka.send(Topics.ORCH_REPLY_INVENTORY,
new InventoryReply(command.commandId(), command.orderId(), reserved, reason));
}
InventoryOrchestrationHandler.java
claim() is the whole pattern, and it’s deliberately not clever:
@Transactional(propagation = Propagation.REQUIRES_NEW)
public boolean claim(String commandId) {
try {
processed.save(new ProcessedCommand(commandId));
return true;
} catch (DataIntegrityViolationException alreadyClaimed) {
return false;
}
}
The UNIQUE constraint on commandId — not an if statement in
application code — is what makes this safe if two threads ever race on the exact same
command: at most one insert can win. Publish the same command, same id, twice in a row:
ProcessedCommand rows for commandId f07ec35c-853c-4405-a024-b62d304dfabb: 1
Stock remaining for GADGET-DUP after two deliveries of the same command: 7
InventoryReply messages actually published: 1
reply: InventoryReply[commandId=f07ec35c-853c-4405-a024-b62d304dfabb, orderId=999, success=true, reason=null]
Full transcript: output/03-idempotent-consumer-duplicate.txt
This guards the command, not the whole saga. If a listener does two things — decrement stock, then also update a shipping estimate — a crash between them and a retry would still only re-run the whole method once more, because the guard is keyed on the command, not on each individual side effect inside it. A listener with multiple side effects still needs each of them to tolerate being retried together, or needs its own finer-grained idempotency keys.
- Delivery guarantees, from the module this saga builds on: kafka-basics companion repository
Choreography vs orchestration: a failure matrix
Both versions of this saga produce identical outcomes for identical inputs — every transcript above proves that. What actually differs between them only shows up once something goes wrong, or once the saga has to grow:
| Situation | Choreography | Orchestration |
|---|---|---|
| A participant never replies (crashes mid-step) | Every other participant is still waiting on an event that will never arrive; nothing times out unless each listener separately implements one | The orchestrator is the one place a timeout belongs, and it already knows exactly which step is outstanding |
| Adding a fourth participant to the saga | New listener, subscribed to an existing event — no existing class changes | The orchestrator’s code changes to know about the new step; every existing participant stays untouched |
| “Why was this order cancelled?” | Reconstructed by reading every listener’s subscriptions in turn, or from a trace spanning several services | Answered by one class’s logic, in order, in one file |
| A duplicate/redelivered message | Exactly as dangerous as in orchestration — the idempotent-consumer pattern above applies to either style equally | Same |
| Single point of failure | None of these classes is more critical than another — but there’s also no one place to add saga-wide monitoring | The orchestrator becomes the thing that must stay up and be deployed carefully; it’s also the one dashboard you’d build first |
Richardson’s own framing is the simplest way to decide between them: choreography costs you less up front and avoids a central dependency, right up until the number of participants makes “what happens next” something you can no longer hold in your head. Three participants, as in this post’s demo, is comfortably on the choreography side of that line for most teams; six or seven participants with multiple failure branches is usually where orchestration’s one readable file starts winning.
Should you reach for a saga at all? Only once a business operation genuinely spans more than one service’s own data, with no way to make it one local transaction. If everything the operation touches lives in a single service’s own database, that’s a local transaction, full stop — sagas trade ACID guarantees for reach, and you only need that trade when the reach is unavoidable. Spring Kafka’s own KafkaTransactionManager can’t make this problem go away either: chaining it with a database transaction manager commits the database transaction, then the Kafka transaction, as two separate commits in sequence — a narrower window than no coordination at all, but not the same guarantee a saga’s explicit compensations give you.
Further reading
- saga companion repository — versions, quickstart, and the output/ index
- Transactional Outbox with the Spring Modulith Event Publication Registry — this series’ post on the dual-write problem inside a single service, the saga’s closest relative at smaller scope
- microservices.io — Saga
- Spring Kafka reference — Transactions
No Comments yet!