CQRS and Projection Lag
Why This Matters
Split a write model from a read model and you've quietly split "true" into two versions of the truth that disagree for a while. Almost nobody designs who owns that disagreement window, how wide it's allowed to get, or what happens when it blows past the limit. They just find out at 2am.
The 2am Page
Write succeeds. Confirmation email goes out. Thirty seconds later the ops dashboard still reads zero orders for the hour. On-call gets paged, assumes a write bug, burns an hour tracing the order service — and finds nothing broken. The read model just hasn't caught up. The system is doing exactly what it was built to do. What's missing isn't a fix, it's an answer nobody wrote down: how long is "catching up" allowed to take?
One Entity, Two Masters
Rewind to before the split. One real Order entity, one ORM model, doing two jobs that have nothing to do with each other. The write path checks invariants: can't ship more than what's in stock, can't confirm payment that hasn't cleared, can't skip a status transition. The read path wants denormalized, pre-aggregated numbers for a dashboard: orders per hour, revenue by region, status counts. Same shape, two masters pulling in opposite directions — and the shape was lying to one of them long before anyone noticed.
- ▹Nullable fields the write path needs for staged invariant checks are dead weight to the read path's aggregation query
- ▹A status enum that's correct for transactional integrity now needs an expensive join on every dashboard refresh
- ▹Foreign keys protecting write-side referential integrity are the exact thing tanking read-side query performance at scale
- ▹A cached display_total added for read-side convenience quietly becomes a second source of truth the write path never updates
The Bug Was Never the Code
Each of those gets patched one at a time — an index here, a cached column there, a nullable field nobody remembers the reason for. None of those patches is wrong by itself. The actual bug sits upstream of all of them: pretending the write side and the read side ever wanted the same model in the first place. That pretense is what produces schema compromises nobody owns and migrations nobody can fully reason about, because every migration is trying to satisfy two masters that were never trying to agree.
What CQRS Buys You — And What It Costs
CQRS doesn't add a new capability. It names a split that was already there and gives each side a model shaped for its one job: the write model enforces invariants and doesn't know dashboards exist; the read model is denormalized and fast and doesn't know what a stock reservation is. That's the whole trade. The cost is specific and you don't get to negotiate it away: you lose the ability to ask "is this consistent right now?" for free. One shared table made consistency a side effect. Two models make consistency something you build and watch — a projection pipeline, a lag metric, a failure mode with a name.
Lag Is Now a First-Class Fact
This is the part the adoption doc always skips. Once you split, lag isn't an edge case to deal with later — it's a permanent, load-bearing fact of the system from day one. Someone decides how stale is acceptable: 200ms, five seconds, thirty during a traffic spike. Someone instruments it as a named metric. Someone owns the incident when it blows past that number. Skip all three and on-call does them anyway, implicitly, at 2am, by guessing.
CQRS in the Age of Agents
Here's where it gets sharper for AI-era systems. Your "read model" today is often a vector index, a RAG embedding store, or materialized context an agent reaches through a tool call — get_order_status, get_customer_history. Same write/read split, worse failure mode: a human staring at a stale dashboard can notice the number looks off and wait a beat. An agent calling a stale read-model tool states the wrong answer as fact — to a customer, or to another agent downstream in a pipeline — with full confidence and zero hesitation. In a fleet of agents, one agent's write (place order) and another's read (where's my order) cross this exact lag window constantly. If staleness isn't surfaced as data, it gets surfaced later as a hallucinated-sounding wrong answer instead.
{
"tool": "get_order_status",
"returns": {
"status": "string",
"as_of": "2026-10-05T09:14:03Z",
"projection_lag_seconds": 4,
"max_staleness_sla_seconds": 30
}
}
// System prompt instruction for the agent:
// If projection_lag_seconds exceeds max_staleness_sla_seconds,
// say "this may not reflect the latest update" instead of asserting fact.That as_of field is the SLA made legible to the one consumer that will otherwise treat every read as gospel. Design CQRS for an agent-facing system and the staleness contract has to live in the tool's return shape, not just in a runbook nobody reads mid-incident.
Quick Gut-Check for Day 23
- ▹Does the write invariant actually conflict with the read shape — or are you just dodging a join?
- ▹Can your consumers, human or agent, tolerate the lag window you're about to introduce — and do they even know it exists?
- ▹Who owns the incident when write and read diverge past the SLA — a name, not a team Slack channel?
Close
CQRS doesn't remove the lie a single model was quietly telling one side of your system. It relocates it — out of the schema, where it sat invisible and unowned, into an SLA, where it's explicit and somebody's name is on it. That relocation only pays off if you actually say the SLA out loud. Skip that step and you haven't fixed anything — you've just moved the 2am page from "why is this schema so weird" to "why is the dashboard wrong." Same unowned lie, new address.
Extend your knowledge
- ▹Read Martin Fowler's bliki entry on CQRS — the canonical warning against applying it to every entity instead of the ones with a genuine write/read conflict
- ▹Look at Greg Young's original CQRS talks/writing, where the pattern was named from real event-sourced systems, not as a scaling checkbox
- ▹If you run a RAG or agent-tool system, check whether your retrieval layer exposes a freshness timestamp the agent's prompt can actually reason about
- ▹Instrument projection lag as a named, alertable metric (e.g. projection_lag_seconds) instead of discovering it during an incident
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.