The Saga Pattern
Day 21: Saga Pattern — A Confession Dressed as Architecture
Here's the tell I watch for now: the same two service names keep showing up together across every saga in the system. That's not a coordination gap you patch with better orchestration tooling — it's a modeling mistake wearing distributed-systems clothing. And it matters more than it did five years ago, because agents are starting to write and run these multi-step workflows themselves. Every saga you ship today is a failure mode an agent will inherit tomorrow, not question.
Nobody asked the obvious question
I once sat in a design review where a four-step saga got rubber-stamped in under ten minutes. Order reserves stock. Payment charges the card. Inventory confirms the reservation. Shipping schedules a pickup. If any step fails, four compensating transactions fire in reverse to unwind the mess. Everyone nodded at the state diagram — it looked clean, it looked thorough. Nobody asked the one question that actually mattered: why do Order and Inventory need to be atomically consistent with each other in the first place? Answer that honestly, and the saga never makes it to a Jira ticket.
What a saga actually is
A saga is a chain of local transactions spread across services, where each step carries its own compensating transaction to undo it if something downstream fails — because you've already given up on two-phase commit and distributed locks across service boundaries (that's Day 19 territory). You coordinate it one of two ways. Orchestration puts a central brain in charge — a saga orchestrator like Temporal, AWS Step Functions, or Camunda — calling each service in turn and tracking the state machine. Choreography skips the central brain entirely: each service reacts to events and fires its own (OrderPlaced → InventoryReserved → PaymentCharged), and the saga exists only as the sum of those reactions. Either way, you're trading atomicity for eventual consistency — and the price is that you now owe the system a hand-written undo for every step a database rollback can't give you for free.
Count the repeat offenders
Here's the diagnostic that actually works: pull up every saga running in your system and tally which service pairs keep showing up together. If Order and Inventory appear in three or four different sagas — reserve-and-release, cancel-and-restock, return-and-replenish — that's not a sign you need smarter coordination tooling. That's a sign the two 'services' are really one aggregate that got sliced apart along a deployment diagram instead of a business invariant. A genuine bounded-context split produces sagas that fan out across many different pairs. A boundary mistake produces the same two names, saga after saga, each time in a new costume.
Case: the classic Order/Inventory saga
Walk through this one honestly. The business invariant is simple: never sell what you don't have in stock. That invariant spans exactly two tables — orders and inventory counts. Split those into two services and the invariant doesn't go away; it just becomes impossible to enforce with a single transaction, so you build a saga to approximate it after the fact — reserve stock, attempt payment, release stock on failure. Nine times out of ten, this is one aggregate in Eric Evans' sense: order-placement and stock-decrement sit inside the same consistency boundary, and they got separated because of a deployment chart, a team-ownership chart, or an 'everything is a microservice' mandate — not because the business logic demanded it. Merge Order and Inventory back into one service with a real transaction, and the saga — all four of its compensating transactions — simply stops existing. You can't have a bug in code you deleted.
Steelmanning the counter-argument
Sagas aren't always a confession. Sometimes the business capabilities really are independent. Order and Shipping, for instance — shipping providers, rate contracts, and carrier integrations change on a completely different cadence, owned by a completely different team, than order placement. Or Order and Fraud-Scoring, where the fraud model scales on its own (often calling out to an LLM-based risk classifier under real-time latency pressure) and is owned by a risk team with its own compliance obligations and release cycle. In cases like these, a saga is the honest tool, not a workaround. The test I use: imagine the same three-person team owned both services tomorrow — would they still ship them as two separate deployables? If yes, because the lifecycles, scaling profiles, or team boundaries are real and would survive even under unified ownership, the saga earns its keep. If the honest answer is 'no, we'd just merge them,' the saga is compensating for an org chart, not a business reality.
- ▹Second test: can the invariant survive being eventually consistent for the life of the saga? If 'stock briefly shows available while payment is still processing' is something the business can live with, choreography is fine. If it isn't, what you need is one transactional boundary, not an undo button.
A checklist for design reviews
- ▹How many sagas already touch these same two services? More than one is the bounded-context smell — stop and ask why before you approve a third.
- ▹What's the actual business invariant here — and can it survive being eventually consistent for the life of this saga, or does the business genuinely need it enforced synchronously?
- ▹Who owns the compensating-transaction code, and who gets paged when the undo fails halfway through its own undo?
- ▹If the same small team owned both services, would they still ship them separately? If the answer is no, the saga is a symptom, not a solution.
Why this bites harder now
I've started seeing teams wire Temporal or LangGraph-style workflows underneath an agent so it can chain tool calls across MCP servers — book a flight, charge a card, reserve a hotel — with a compensating step if the last call fails. That's a saga with an LLM sitting where the orchestrator's brain used to be. The question is the same, except now the boilerplate is cheaper than ever to generate: an agent will happily write you a syntactically correct compensating transaction for a service boundary that never should have existed. It won't push back on your architecture. You have to — before the saga ships, not after the 3am page.
Close
Sagas aren't free. Every compensating transaction is a new state machine, a new failure mode, and one more thing an on-call engineer — or an agent acting on their behalf — has to reason about under pressure. Every saga you avoid by fixing the boundary instead is a bug you'll never have to debug at 3am, because it never got the chance to exist.
Extend your knowledge
- ▹Chris Richardson's microservices.io write-up on the Saga pattern — the clearest reference for orchestration vs. choreography trade-offs.
- ▹Pat Helland's essays on distributed transactions ("Life beyond Distributed Transactions") for why 2PC falls apart at scale.
- ▹Eric Evans' Domain-Driven Design — the chapter on the Life Cycle of a Domain Object, where he lays out Aggregates — for how to find the correct consistency boundary before you reach for a saga.
- ▹Temporal's documentation on implementing the Saga pattern — useful if you're now using it to orchestrate multi-step agent/tool-call workflows, not just microservices.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.