Back to blog

Same-pass verification bias

Oct 5, 2026
Series · Day 19
Software Engineering in the AI Era
View all lessons →
Same-pass verification bias

Day 19 — When the Fix and the Test Share One Brain

A green checkmark on CI feels like proof. It isn't — not when the same agent pass wrote both the fix and the test meant to catch it. All you've confirmed is that the agent agrees with itself. The bug can ship twice and the board still looks clean.

The incident

An agent picked up a dedup bug — duplicate events were leaking through a webhook processor. It diagnosed the cause, wrote the fix, wrote a regression test for that fix, and opened the PR. CI went green: new test passed, nothing else broke. We merged it. Four hours later the same duplicates showed up in prod again, from the exact root cause the PR was supposed to kill.

Pull the diff apart

Here's the agent's mental model: duplicates are events sharing an idempotency key that arrive inside the same 60-second window. So the fix truncated timestamps to the minute and compared keys inside that bucket. The new test asserted exactly that — two events, same key, same minute, correctly flagged as a dupe. It never asserted the condition that actually happened in prod: client retries landing seconds apart but straddling a minute boundary — 12:00:59.8 and 12:01:00.2.

python
# the fix (agent's assumption: dupes land in the same minute bucket)
def is_duplicate(event, seen):
    bucket = (event.key, event.timestamp.replace(second=0, microsecond=0))
    return bucket in seen

# the regression test the SAME pass wrote
def test_dedup_same_minute():
    e1 = make_event(key="abc", ts="12:00:05")
    e2 = make_event(key="abc", ts="12:00:45")
    assert is_duplicate(e2, seen={bucket_of(e1)})
    # never tests: ts="12:00:59.8" vs ts="12:01:00.2" — the actual prod case

That test is green for the wrong reason. It checks the consequence of a flawed assumption, not the condition that actually caused the incident. It isn't a weak test, either — it's a perfectly faithful mirror of the same broken model that produced the fix in the first place.

Why this isn't just 'a skipped test'

  • ▹Manual TDD, done properly, is two separate acts of reasoning: someone writes a test from the bug report and the prod symptoms, then someone writes the fix to satisfy it. Even solo, there's a context switch between 'what should be true' and 'how do I make that true'
  • ▹That switch is exactly where contradictions surface. If the fix doesn't actually address the failure mode you tested for, the test stays red, and you notice
  • ▹One-shot agent generation collapses both acts into a single inference over a single mental model of the bug. The fix and the test come from the same internal representation, so the test is structurally incapable of falsifying the fix's core assumption
  • ▹It's the same flaw as a student grading their own exam with an answer key they wrote while taking it — what gets measured is consistency, not correctness
  • ▹And it's invisible in CI, because the artifact looks exactly like real TDD: new code, new test, green run. The failure isn't that a test is missing — it's what that test encodes

Why 'ask the agent to write better tests' doesn't fix it

Telling the agent 'also check edge cases' or 'write more thorough tests' just reruns the same blind spot with extra words attached. The agent doesn't know its model of the bug is wrong — if it did, it would have written a different fix. Asking it to grade its own work more rigorously still routes through the same diagnosis. What you need is a second source of reasoning that never inherited the first pass's assumption — not a longer prompt to the same one.

The fix we adopted at PhoenixDX

  • ▹Regression tests for an agent-written fix now come from a second, separate pass — another agent invocation, or a human — shown only the bug report and the prod symptoms
  • ▹That second pass never sees the fix's diff, the PR description, or the reasoning trace. If it can read the fix, it can anchor on the fix's assumptions and mirror them right back
  • ▹The test from pass two gets written and reviewed before anyone checks it against the fix. If it happens to pass anyway, that's useful signal — not proof
  • ▹In practice: two CI jobs, two agent sessions, disjoint context. Session A gets the bug ticket and writes diagnosis plus fix. Session B gets the same ticket plus the prod logs and traces and writes the test — no access to session A's branch until grading time

Day 20 teaser

Obvious next question: if pass two is also an agent, what stops it from converging on the exact same wrong model on its own? Or a third failure mode — the checker turning out lazier or more literal than a human reviewer would ever be? Tomorrow: how we checked the checker.

Flashcards
Check yourself

Extend your knowledge

  • ▹Audit a recent agent-authored PR in your own repo. Find the regression test and ask: does it assert the prod symptom, or just the fix's internal logic?
  • ▹Try the two-pass split on one bug this week — one agent session for diagnosis and fix, a second with only the ticket and logs for the test — and see what the second pass catches that the first missed
  • ▹Read Kent Beck on TDD's red-green-refactor cycle to see why the two-reasoning-pass structure matters, then map it onto how you run agent workflows
  • ▹Look into the 'oracle problem' in software testing literature. This incident is just a specific, AI-era instance of a decades-old challenge
Test yourself on this lesson →

Discussion

Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.

Ask me anything about “Same-pass verification bias” — trade-offs, decisions, or the story behind it.