Back to blog
Series · Day 18
Software Engineering in the AI Era
View all lessons →
Debugging agent loops

Day 18: The Harness Lied, and the Agent Was Right to Believe It

An agent stuck looping on a two-line fix looks like a broken brain. Nine times out of ten it isn't — the model is reasoning just fine, on a lie someone else fed it.

The incident

One of our PhoenixDX agents got stuck doing the same thing over and over — same diagnosis, same fix, same tool call, retry count climbing past ten. The task behind it was nothing: a config update gated behind a deploy check, the kind of thing we expected to close in two or three turns. Nobody on the team could say why it wouldn't stop. Read the transcript cold and it looks exactly like an agent that's lost the thread.

First instinct, and what we tried

We did what everyone does first: assumed the model's reasoning was broken. Swapped in a bigger model. Rewrote the prompt to spell out the stop conditions in plain language. Gave it a bigger retry budget on the theory it just needed more room to find its own way out. None of it worked — the loop count went up, not down, because all we'd actually done was tell it to try harder at the wrong thing.

  • ▹Swapping in a bigger model changed nothing — there was no miscalculation to fix
  • ▹Spelling out stop conditions got ignored, because by the agent's own state, none of them were met
  • ▹Raising the retry budget just stretched the loop out, didn't shorten it
  • ▹Every single attempt was the logical move, given what the agent believed was true

The actual discovery

We stopped tuning the agent and went turn-by-turn through the transcript against the raw tool I/O — not the agent's summary of what happened, the actual bytes. The tool call right before the loop started had hit a 429 from a downstream API. The harness's retry wrapper caught it, retried silently, burned through its own retry budget, and then — instead of surfacing the failure — handed the agent back a generic 'success.' So the agent was told: done, it worked. It moved on, checked the goal state, found nothing had changed, and drew the only conclusion that made sense: its fix hadn't landed yet. So it tried again. And again.

Why this matters mechanically

Here's the thing: the agent's next move was the correct response to what it was shown. Told 'the last action succeeded' and seeing 'the goal condition is still false,' the only sane read is 'something else must be wrong, try a different angle.' That's not a broken decision function. That's the decision function doing exactly its job on corrupted input. The bug was never in the model's reasoning — it was in the layer sitting between the model and the world. The harness's error handling was lying about ground truth, and the agent had no way to catch the lie, because it has no channel to reality except what the harness reports.

The generalizable check

Before you touch the prompt or swap the model, ask one question: does the agent's context match reality, or does it match what the harness told it? This matters more for agents than it ever did for regular software, because an agent has no senses of its own — it never hits the API directly, never queries the database itself, never watches a log stream with its own eyes. Everything it 'knows' about the world passes through a tool wrapper, an SDK, a retry policy, an error handler — something a person wrote, often you. Anywhere that layer coerces, swallows, or paraphrases an error into something agent-friendly, the agent's model of the world can quietly drift from ground truth. Once it drifts, the behavior you're seeing stops being explainable by the prompt or the weights.

A fast diagnostic habit

When you spot a loop, don't touch the prompt first. Find the last tool call before the loop started. Put its claimed result next to its raw result — not the agent's interpretation, the actual bytes the tool handed back: status codes, headers, whatever your wrapper logged before it got transformed. Check the claim against ground truth — did the file actually change, did the deploy actually run, did the API actually return 200. If the claim and the ground truth disagree, you've found a harness bug, and no prompt engineering on earth fixes that. If they agree, now you've got an actual reasoning failure worth debugging at the model layer.

  • ▹Find the last tool call right before the repeated behavior kicked in
  • ▹Pull the raw tool response — status code, payload, logs — not the agent's summary of it
  • ▹Verify independently against ground truth: did the real-world effect actually happen
  • ▹Only touch the prompt or the model once you've confirmed the harness was telling the truth

Where this fits in the arc

Debugging agents is systems debugging with an LLM sitting in the loop — it's not some new discipline that needs its own mythology. The same muscle you'd use on a distributed system retrying against a flaky dependency applies here, unchanged: check the boundary, trust nothing that's been summarized, verify against ground truth. Tomorrow builds on this directly — treating the harness itself as a first-class suspect, not an assumed-correct platform, every time an agent's behavior looks irrational.

Flashcards
Check yourself

Extend your knowledge

  • ▹Go audit your own tool wrappers: grep for retry/catch blocks that return a fixed success payload on any non-happy path — that's this bug's exact shape
  • ▹Next time an agent loops, log the raw tool responses — status codes, not just parsed text — next to the transcript so you can diff claim against reality fast
  • ▹Look at how well-designed agent SDKs structure tool-call results: status, raw payload, and parsed summary kept separate, so an error has to surface instead of getting coerced into a fake success
  • ▹Tomorrow: treating the harness as a first-class debugging suspect, not an assumed-correct platform
Test yourself on this lesson →

Discussion

Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.

Ask me anything about “Debugging agent loops” — trade-offs, decisions, or the story behind it.