Back to blog
Series · Day 18
Solution Architecture in 30 Days
View all lessons →
Saga Compensation Taxonomy

Day 18 — Compensation Is Not Undo: The Three-Bucket Reversibility Taxonomy

Every saga diagram I've ever reviewed draws the same lie: a clean 'compensate' arrow pointing back to zero. It's lying about at least one step, in almost every real workflow you'll ever ship. Skip the question of which steps can genuinely be undone before you sequence them, and you'll end up with a rollback path that tries to reverse something that was never reversible to begin with.

What's the compensation for 'shipped'?

Take the saga every tutorial has drawn at least once: reserve inventory → charge card → ship package. Say step three fails somewhere downstream — a damaged-item flag trips, a fraud hold fires late, doesn't matter which. Your diagram says run the compensations in reverse order: un-ship, refund, release the reservation. Un-charge is awkward but workable — you issue a refund. Un-reserve is trivial — you release the lock. Un-ship isn't a thing. The truck already left. There's no API call that teleports a package back into the warehouse. The diagram is pointing an arrow at an action that doesn't exist.

'Rollback' is database vocabulary, and it snuck in uninvited

The saga pattern borrows its mental model from database transactions, where ROLLBACK is guaranteed, atomic, and free — the database never actually committed the write. A compensating transaction in a saga is nothing like that. It's a brand-new, forward-moving transaction that happens to point in a corrective direction. That distinction isn't a footnote, it's the whole design problem: a new transaction inherits every property of any other transaction. It can fail. It can partially apply. It can take time, cost money, or — as with 'un-ship' — simply not exist as a definable operation at all. Calling it 'compensation' instead of 'rollback' isn't pedantry. It's the entire design problem, compressed into one word.

The three-bucket taxonomy

  • ▹Reversible — a true undo exists and costs next to nothing: release a lock, cancel a pending reservation, delete a draft record. These are the only steps where 'rollback' is actually an honest word to use.
  • ▹Semi-reversible — you can neutralize the effect, but it costs something: money, trust, a manual step. Refund a charge minus the processing fee. Issue store credit instead of cash. Apologize and comp the customer. The state gets corrected — just not for free, and not perfectly.
  • ▹Terminal / irreversible — nothing undoes it, you can only mitigate the fallout afterward: an email that's already sent, a package that's already shipped, a webhook already fired to a partner, an agent's external API call (a tweet posted, a ticket filed, an invoice paid). Once it fires, you're in incident response, not rollback.

What this taxonomy tells you about ordering

Once every step has a bucket, sequencing stops being a judgment call and becomes arithmetic. Terminal steps go last, always. Reversible steps can run early, because backing them out is cheap. Semi-reversible steps belong in the middle, because backing them out is possible, just lossy. Terminal steps have to be the final gate — the one thing you do only after everything upstream has already succeeded — because the moment one fires, you're no longer running a saga, you're running an incident. That's why 'just reorder the saga' sounds hard but turns out to be mechanical: the taxonomy already told you the order. You just have to read it off.

Worked example: redesigning 'reserve → charge → ship'

The textbook version sequences by narrative convenience — reserve, then charge, then ship — because that's the order a checkout flow reads naturally to a human. The taxonomy version sequences by what you can afford to be wrong about last.

text
Step              Bucket            Compensation
----------------  ----------------  --------------------------------
reserve_inventory reversible        release the reservation
charge_card       semi-reversible   refund (minus processor fee)
ship_package      terminal          none — only post-hoc mitigation
                                    (recall request, customer support,
                                     replacement shipment)

Notice the order didn't actually move here — reserve → charge → ship happened to already be terminal-last, by luck. The taxonomy earns its keep the moment someone adds a 'helpful' step like send_confirmation_email or notify_partner_webhook. Teams bolt that straight onto the happy path right after charge_card, because that's where it reads naturally in the code. The taxonomy forces a different question: is this step terminal? Yes — it's an email, you can't unsend it. So it moves to the end, after ship_package succeeds, not before.

Why this matters more in agentic pipelines

In my ASE research and in the day-to-day at PhoenixDX, the same pattern shows up over and over: LLM-agent pipelines pack a much higher density of terminal steps than a typical microservice saga. A single agent turn can send a Slack message, write an embedding into a vector store, trigger a downstream automation, call a paid third-party API, or kick off another agent — and every one of those is terminal the instant it fires. In a normal microservice saga, you can usually count the truly irreversible steps on one hand — often it's just the one. In an agent pipeline, most steps are terminal, because 'take an action in the world' is the entire point of giving an agent tools in the first place. Most agent frameworks today hand you retries, timeouts, structured outputs — none of them force you to declare which tool calls are reversible before you wire up a multi-step plan. That gap is yours to close. Audit your tool list the same way you'd audit saga steps.

The actionable takeaway

Before you draw a single saga arrow — or wire up a single multi-step agent plan — write a three-column list: step, bucket (reversible / semi-reversible / terminal), compensation (or 'none, mitigation only'). That list is the actual design artifact. The sequence diagram, the state machine, the LangGraph or Temporal workflow definition — all of it is just a rendering of the list, produced after the classification is done, never before.

Flashcards
Check yourself

Extend your knowledge

  • ▹Read Garcia-Molina and Salem's original 1987 'Sagas' paper — it's where the compensation terminology comes from, written for long-lived database transactions, not microservices.
  • ▹Look at how Temporal documents the saga pattern and how AWS's prescriptive guidance describes saga compensation on Step Functions — both treat compensating actions as ordinary, fallible activities, not guaranteed rollbacks.
  • ▹Audit your own agent's tool definitions: tag each one reversible, semi-reversible, or terminal before wiring it into a multi-step plan or workflow graph.
  • ▹Compare this taxonomy against the outbox pattern — it's a related answer to the same problem: combining terminal external effects with eventual consistency.
Test yourself on this lesson →

Discussion

Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.

Ask me anything about “Saga Compensation Taxonomy” — trade-offs, decisions, or the story behind it.