Back to blog
Series · Day 18
Multi-Agent Systems in 30 Days
View all lessons →
Designing the Observe Step

Day 18 — Your Loop Isn't Slow Because of the LLM. It's Slow Because Nobody Designed 'Observe.'

Here's a number that'll ruin your morning: check how many of your agent's iterations today actually shipped something, versus how many just sat there waiting for someone to glance at them. Most teams time their agent loop by generation speed — tokens per second, iterations per day. But the number that decides how much actually gets out the door is how fast you can trust what the thing just produced. Almost nobody designs for that on purpose. It just happens to them.

The throughput math nobody runs

Pull the logs on a 'working' plan-act-observe loop and you'll usually see something like this: 40 iterations in a day, but only 6 of them actually moved state forward — merged, deployed, bumped to the next stage. The other 34 were the agent generating, re-generating, and waiting. Not because the model was slow. Because a human had to look at the output before anything downstream could move, and that human had six other tabs open.

Iterations-per-day is a vanity metric. Advances-per-day is the real one, and it's gated by whoever — or whatever — gets to say 'yes, this is good enough to continue.' If that gate is a person eyeballing a Slack thread between meetings, your agentic pipeline runs at the speed of your busiest reviewer's calendar. Not your inference endpoint.

Name the pattern: generator bolted to reviewer

This isn't a hypothetical — it's what our early rollout at PhoenixDX actually looked like. We had a generator that was fast, cheap, and trivially parallel — spin up twenty of them for pennies — bolted straight onto a reviewer that was slow, serial, and human. The generator could crank out a week's worth of candidate work before lunch. The reviewer got through maybe three or four of those a day, because reviewing agent output properly takes real attention, and humans don't parallelize no matter how much coffee is involved.

The mismatch never shows up as an error. It shows up as a backlog that quietly fattens, a team that feels busy, and an 'agent activity' graph that looks great right up until someone asks what actually got merged this week.

Why this is invisible: 'observe' sounds automatic

Every agent-loop diagram draws the same three boxes: plan, act, observe. Observe sounds passive, like the agent just notices what happened and moves on. It isn't passive. Observe means verify — did the action actually do what we wanted, is it correct, is it safe to build on top of. That's a real engineering step with a real cost. And in nearly every team I've watched, including my own early on, that cost gets quietly outsourced to 'a human will check it.'

Nobody writes that down as a decision. It just happens, because in week one, with one agent and low volume, having a human check the output feels completely normal — it's what you'd do anyway. The cracks only show once you scale the generator and forget to scale the verifier, which agentic tooling makes shockingly easy to do by accident.

The tell

Here's the diagnostic: a loop with no real verification step has exactly one speed — generation speed. What you actually have is a queue wearing a trenchcoat, pretending to be a pipeline. You can tell because the moment you ask 'what happens if I 10x the generator,' the honest answer is 'the backlog 10x's, nothing else changes.' A real verification step, on the other hand, is something you could scale too — more compute, more parallel checks, a sharper model. A queue scales by hiring more humans, and you can't do that fast, and you don't actually want to.

What a designed verification step looks like

  • ▹A cheap, fast, machine-checkable signal that runs on every single iteration — unit tests, schema validation, type checks, a linter, a contract test against a real API
  • ▹A second agent acting as a narrow critic — not 'is this good' in the abstract, but one specific, checkable claim: does this diff compile, does this output match the schema, does this answer actually cite a real passage from the source doc
  • ▹Tiered gates — cheap, fast checks run on every output and reject obvious failures on the spot; expensive, slow checks (including humans) only ever see what already cleared the cheap gate
  • ▹Human review demoted to spot-check — sample a percentage of passed outputs for quality drift, instead of gatekeeping every single one before it's allowed to move
  • ▹A clear answer for what it means when the fast check passes but the output is still bad — so you're not flying blind between spot-checks

Where this goes next

'Add a verifier' is the easy half of the sentence. The hard half — where Day 19 picks up — is that not every check is fast enough or trustworthy enough to actually close the loop. A test suite that takes 20 minutes doesn't close a loop that's meant to iterate in seconds. A critic agent that rubber-stamps whatever the generator hands it — because it trained on similar data, or because its prompt is too soft — gives you false confidence dressed up as throughput. What makes a verifier actually good — fast enough, independent enough, hard enough to game — is its own design problem, and it's the next thing worth getting precise about.

The diagnostic question for this week

Stop asking how good your model is. Ask this instead: what's the slowest mandatory step between two agent actions in my loop right now — and did I actually design that step, or did it just happen because it felt natural back in week one? If you can't name it, that's your answer.

Flashcards
Check yourself

Extend your knowledge

  • ▹Pull your own loop's logs this week and compute iterations-per-day vs. advances-per-day — the gap is your real bottleneck, not a guess
  • ▹Look into 'LLM-as-judge' failure modes (sycophancy, reward hacking, position bias) before you lean on a critic agent as your fast verifier — Day 19 builds on this
  • ▹Compare it to how mature CI/CD teams solved the exact same shape of problem with trunk-based development and fast test suites — 'cheap check on every commit, humans sample the rest' predates agents by decades
  • ▹Audit one agentic workflow you run today and write down, in plain words, what the slowest mandatory step between two actions is — and whether anyone actually chose it
Test yourself on this lesson →

Discussion

Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.

Ask me anything about “Designing the Observe Step” — trade-offs, decisions, or the story behind it.