Back to blog

The Three-Correction Rule for Debugging with AI Assistants

Oct 5, 2026
Series · Day 3
Using AI Day to Day
View all lessons →
The Three-Correction Rule for Debugging with AI Assistants

Day 3: Why Judgment Is the Scarce Skill Now

The better the model gets, the more its wrong answers look exactly like its right ones — same confident tone, same clean diff, same ready-to-ship feel. If you can't tell the difference before you hit run, the model's biggest strength is now working against you.

Three sessions, three fixes, three wrongs

Same bug. Three chat sessions. Three days. Orders were occasionally getting processed twice, and every session went the same way — paste the error, get a fix, ship it, close the tab feeling done.

Session one: the AI suggested adding an idempotency key to the request. Already there — a dead end. Session two produced the fix that actually looked right enough to ship:

diff
- await db.orders.insert(order)
+ try {
+   await db.orders.insert(order)
+ } catch (err) {
+   if (err.code === '23505') {
+     // duplicate key — order already processed, safe to ignore
+     return
+   }
+   throw err
+ }

It reads like defensive programming — catch the duplicate-key error, assume it's a legit dupe, move on. It shipped. Two days later, a new bug report landed: customers whose orders silently stopped updating after the first write. Turns out the catch block wasn't catching real duplicates at all. It was swallowing a dedupe-key collision caused by a field rename in a refactor from the week before, and in the process it buried every symptom that would've pointed straight at the real cause. Session three's fix: a frontend debounce on the submit button. Plausible. Also wrong.

The turn

A teammate picked up the thread. She ignored all three fixes and didn't open a chat window at all. First she reproduced the duplicate herself, replaying one real webhook payload locally. Then she read the full stack trace instead of pattern-matching the top line — and one frame stopped her: the dedupe lookup was keyed on `orderRef`, but the webhook payload field had been renamed to `order_ref` in that refactor. Every lookup missed. Every insert looked "new." The database's unique constraint was the only thing catching it at all. She confirmed it with `git bisect` against the commit that renamed the field. Ten minutes. Zero prompts.

What actually happened in the first three sessions

Nobody here was being lazy. Each session felt like forward motion — a new message, a new diff, a fresh reason to believe it was handled. But look closer and what was actually happening was rewording the same under-specified problem and repasting it, not restating it. The error message changed each time. The ask underneath never did. Prompting feels like progress because something happens every single time — a plausible diff, a confident explanation, a fix that compiles cleanly. Reproducing a bug by hand doesn't give you that hit. For five minutes, nothing "happens," which makes it feel slower. It isn't slower. It was the only one of the four moves that was ever going to find a field-rename bug instead of a symptom to paper over.

The rule this earns

If you've corrected the same prompt three times — same bug, new framing each attempt — the bug isn't in the code anymore. It's in your understanding of the code. No fourth prompt, however cleverly worded, fixes a gap in your own mental model of the system. Only reading the actual stack trace, the actual diff, or the actual commit history closes that gap. Here's the one tripwire worth keeping from this whole lesson: count corrections, not minutes. Three is the line. After that, stop prompting and go read something real.

It's not just a debugging quirk

The same failure mode shows up anywhere you let a plausible AI output stand in for a question you never actually answered yourself —

  • ▹Writing — the draft reads fine, but you never pinned down who actually reads this and what they do right after. Clean prose, zero check against a real reader.
  • ▹Design — the schema looks tidy, but you never asked where this data actually lives or what breaks first under load. Clean diagram, zero check against real failure modes.
  • ▹Debugging — the diff compiles and looks defensive, but you never reproduced the failure yourself, so you don't actually know what correct behavior looks like. Clean fix, zero check against the real system.

Same shape, every time: the AI's answer is coherent and well-formed, and coherence was never the thing you needed checked.

The paradox

Strong models made plausible answers free — cheap, fast, available on every single retry, whether you're debugging, writing, or designing. That's a genuine gain. But it quietly moved the bottleneck. The question is no longer "can I generate an answer." It's "can I tell, before I've even seen the answer, what a correct one would look like." That's judgment, and you don't train it by sending a fourth prompt. You train it by reproducing the bug yourself, reading the actual error, building an internal model of correct — the unglamorous basics, repeated enough times that you develop a reflex for when something's off, even before you can say exactly why.

Connect back to Day 1 and Day 2

Reproduce it. Read the error. Know what correct looks like. Boring words — which is exactly why they're the first thing to get skipped under deadline pressure. They don't look like progress the way a fresh diff does. The cost doesn't show up today. It shows up six months from now, in an engineer who ships fixes that happen to work but can't explain why, because every one of those boring checks got quietly outsourced to a model that was never actually asked to do them. It was only ever asked to sound confident.

Flashcards
Check yourself

Extend your knowledge

  • ▹Practice the tripwire this week: the next time you correct the same prompt twice, stop before the third attempt and go reproduce the bug by hand instead.
  • ▹Read David Agans' 'Debugging: The 9 Indispensable Rules' — rule #1, 'Understand the system,' is the pre-AI version of exactly this lesson.
  • ▹Revisit Day 1 and Day 2 of this series — this lesson is what happens when the basics from those two get skipped under time pressure.
  • ▹Next time an AI fix 'works,' spend two extra minutes asking: what would I have expected to see if this were actually correct — and did I check that, or just that it compiled?
Test yourself on this lesson →

Discussion

Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.

Ask me anything about “The Three-Correction Rule for Debugging with AI Assistants” — trade-offs, decisions, or the story behind it.