Event Sourcing
Day 22: Event Sourcing — The Replay-vs-Remember Bug Every Agent Stack Has
Your agent says something false — a config value nobody set, a decision the user never made — and you go to find out why. You can't. The context's been summarized four times since, and a summary, once written, erases the thing it summarized. You open summary #3. Plausible. You open summary #2. Also plausible. The error could've slipped in at any of those four collapse points, and there's no path back to ground truth. You're not debugging anymore. You're guessing which round of lossy compression ate the fact you needed. Event sourcing solved exactly this, for databases, long before anyone was prompting an LLM: keep the raw log, derive state by folding over it, snapshot only when you need the speed.
The corrupted-memory moment
Walk through it slowly. The agent's been running two hours, across dozens of tool calls. It states something false with full confidence. You go looking for where it went wrong. By now the context window has been summarized four separate times. You crack open summary #3 — looks fine. Summary #2 — also fine. The bug could have entered at any of those four points, and once a summary overwrites what it was built from, there's no function that hands you the original turns back. You've stopped debugging and started guessing which lossy compression pass swallowed the truth.
Name the bug class: summarization has no inverse
This isn't a prompting problem. It's a data-modeling problem. Summarization is lossy compression, full stop — and most agent frameworks treat the summary as the new source of truth, then throw away what it came from. Once that happens, no function exists that takes the summary and hands you back the original turns. You swapped a replayable history for one unverifiable opinion about that history, and then kept summarizing the opinion.
Today's term, precisely: event sourcing
Event sourcing is a storage pattern, and it has four parts you don't get to skip:
- ▹Store the append-only log of what happened — every event, in order, never edited, never deleted.
- ▹Derive current state by folding over the events — replay them in sequence to compute where things stand right now.
- ▹Snapshot occasionally for performance — a cached fold result, not a replacement for the log.
- ▹Never overwrite history. The log is the only source of truth; everything else is just a view computed from it.
The reframe: your agent's transcript is already an event stream
A turn-by-turn agent log — user message, assistant message, tool call, tool result, user message, tool call — is structurally the same thing as an event stream in any CQRS/ES system. You just haven't named it that yet. The question was never whether to adopt event sourcing. It's whether you keep treating that log as disposable scratch paper, or finally treat it as the durable record it already behaves like.
[
{"seq": 1, "type": "user_message", "content": "Deploy the staging config to prod"},
{"seq": 2, "type": "tool_call", "tool": "read_config", "args": {"env": "staging"}},
{"seq": 3, "type": "tool_result", "content": {"replicas": 3, "region": "us-east-1"}},
{"seq": 4, "type": "assistant_message", "content": "Staging has 3 replicas in us-east-1. Deploying now."},
{"seq": 5, "type": "tool_call", "tool": "deploy", "args": {"replicas": 3, "region": "us-east-1"}}
]
// This is your event log. The context window you hand the model
// next turn is just one projection folded from these events.The Rosetta stone
- ▹Context window → a projection: one derived view of the log, not the log itself.
- ▹Summarization → snapshotting: a cached fold at a point in time, meant to be rebuilt — not the only copy that exists.
- ▹"Why did the agent do that?" → replay: fold the raw events from that point forward and watch the real sequence happen, instead of trusting someone's compressed guess at it.
- ▹Context compaction or truncation → a different projection: same log, folded with a different window or filter.
Why this distinction is load-bearing, not academic
A snapshot, used correctly, is just a cache — wrong or stale, you throw it out and recompute from the log, zero data lost. A summary, built the way most agent stacks build it today, is the only copy that exists. Teams overwrite the transcript with the summary to save tokens, to save storage, and in doing so turn a disposable cache into the sole record. Lose it, botch it, let the model hallucinate while writing it — the truth isn't degraded, it's gone. That's the whole gap between event sourcing done right and summarization done the way most frameworks ship it.
The design rule to take away
Keep the raw event log durable and separate from whatever compressed view you feed the model. Log every message, tool call, and tool result immutably — disk, database, object storage, doesn't matter which — before you compress anything. Treat every context-window assembly, every summary, every truncation as a disposable, regenerable projection computed from that log. Projection wrong? Rebuild it. Log gone? Nothing saves you.
When this is overkill
- ▹Short-lived, single-session agents that will never need an audit, a debug session, or re-derivation — the storage and engineering overhead just isn't worth it.
- ▹Prototypes where you're still figuring out what state even matters — modeling events too early slows down the exact iteration you need.
- ▹Don't confuse 'store everything forever' with event sourcing. The pattern is about immutability and replay, not hoarding — you still need retention and compaction policies, just applied to snapshots, never to the integrity of the log.
The throughline
Event sourcing predates agents by decades — it's how banks reconstruct account balances, how Kafka-based systems recover state after a crash. It matters now because every agent fleet you ship is quietly accumulating the same replay-vs-remember problem these systems already solved. Start treating your transcript as the log it already is, and "why did the agent do that" stops being forensics. It becomes a query.
Extend your knowledge
- ▹Read Martin Fowler's 'Event Sourcing' article — a clean definition of the pattern, written for databases, that ports directly onto agent transcripts.
- ▹Watch or read Greg Young's CQRS/ES talks — the foundational replay-vs-snapshot thinking this lesson's framing builds on.
- ▹If your agent framework ships built-in checkpointing (LangGraph and similar), check whether its checkpoints are full event logs or already-collapsed state snapshots. That answer tells you exactly how much forensic recovery you actually have.
- ▹Try it hands-on: log every message, tool call, and tool result from one agent run as immutable JSON events, then write your summarizer as a separate function that folds over that log. You'll notice you can now generate several different projections from the same unchanged source.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.