AI pairing and debugging apprenticeship
Why this matters
A junior can hand you the right fix and still be unable to tell you how they got there. Used to be, that gap surfaced in a week. Now it hides for months, because the fix keeps being correct.
The red flag that stopped being a red flag
A junior on my team fixed a race condition in a queue consumer last quarter. Fix was correct, shipped clean, tests passed. I asked him to walk me through how he found it. 'I pasted the stack trace and the file into the AI, it said race condition on the shared connection pool, I applied the fix it suggested.' That was the whole answer. No theory he'd ruled out first, no log line he didn't trust, no moment where he thought it was something else and changed course. Five years ago that answer gets him sent back to re-debug it properly — it means he got lucky with the first guess. Today it's just Tuesday. The fix was right. Nothing on the dashboard flags it.
Name the mechanism: apprenticeship was never about the fix
The old loop — junior sits next to senior, watches them debug live — never really taught fixes. Fixes are cheap, you can look one up. What it taught was a search process: where to look first, which signals to trust, which hypotheses to rule out before committing to one, how to spot a red herring before it burns an afternoon. That's the actual skill, and it's slow to build precisely because you build it by watching someone else get it wrong out loud, not by seeing the one answer that turned out right.
That belt is still running. It just got rerouted. The junior pairs with a model instead of a senior now, and the model jumps straight to the hypothesis that tests out — it doesn't narrate the three it quietly discarded, because nobody asked it to. Same belt, different cargo: it's carrying answers instead of search process.
Why this is easy to miss
Every metric you'd normally lean on to catch this looks fine, or better. Velocity's up. Code review is clean — the fix is often tighter than what this same junior would've shipped unsupervised two years ago. PRs read competently. None of that tells you whether the junior understood the failure mode well enough to reproduce the diagnosis cold, or just pasted an error and shipped whatever came back. The erosion hides precisely because the tool is good. A bad tool would've outed the gap for you by now.
The concrete failure mode
It shows up exactly when you can't afford it: senior's on PTO or buried in another incident, and the AI is either down (outage, rate limit, air-gapped prod) or confidently wrong — hallucinating a plausible root cause for a failure mode it hasn't actually seen enough of. The junior has no fallback instinct to reach for. Not 'struggles a bit more than usual' — no entry point at all, because they were never the one generating hypotheses. They were the one checking a hypothesis someone, or something, else generated.
- ▹It's a missing dependency, not a skill gap — you can't improvise a search instinct under incident pressure if you've never built one.
- ▹It bites harder in on-call than in day-to-day coding, because incidents are exactly when the senior is least available and the AI's training data is thinnest — novel failure modes, weird interactions between your specific services.
- ▹It compounds across generations: a junior who never built the instinct can't teach it to the junior after them, so the gap doesn't hold steady on a team, it widens each cycle unless someone deliberately steps in.
The prescription: a 5-minute narration checkpoint
Before the junior shows you the fix or the merged PR, have them narrate the path out loud for five minutes — what they thought it was first, what ruled that out, what made them change direction, which signal they actually trusted at the end. You're not grading the fix; you already know it's probably right, that's not the question. You're listening for branches — a search with dead ends in it, not a straight line.
If what comes back is a clean, linear paraphrase of the AI's own explanation — no false starts, no 'I thought it was X until I saw Y' — that's your tell. It means the AI did the searching and the junior did the typing. Don't let it slide. Ask the one question the AI never answered: 'what else could've caused that same symptom, and how would you have told them apart?' That question can't be lifted from a chat transcript — it forces actual reasoning about the failure space, not recall of an explanation.
- ▹Run this on a sample of fixes, not every one — you're calibrating, not surveilling.
- ▹Ask it right after the fix, while the path's still fresh, not three days later in a retro.
- ▹The question that catches paraphrasing: 'what would you have tried next if that first fix hadn't worked?' — a transcript doesn't usually hand you that answer pre-packaged.
If you're the junior: the deliberate-rep move
Once a week, pick one bug and debug it with the AI pair switched off entirely. No pasting the stack trace, no autocomplete-shaped hypotheses. Logs, print statements, bisecting, reading the actual code path. It'll feel slower and less efficient, because it is, in that moment. Reframe it anyway: you're not burning a day of output, you're doing a rep for the instinct the AI-assisted loop doesn't build for you by default. Treat it like conditioning work an athlete does that never shows up on the scoreboard that week — not optional just because it's not immediately productive.
The self-check for everyone past junior
This isn't only a junior problem. Name one skill you've quietly stopped practicing because the tool now does it for you — writing a regex from scratch, remembering a CLI flag, reasoning through a SQL query plan, drafting an architecture doc before generating one. Then ask yourself: was that a decision, or just drift? Delegating a skill on purpose, because you've decided it's no longer load-bearing for your growth, is fine. Losing it by accident, because the tool was sitting right there and you stopped noticing you'd stopped reaching for the muscle, is the same erosion this whole lesson is about — just further up the ladder.
Extend your knowledge
- ▹Try the 5-minute narration checkpoint on your next PR review and notice whether the explanation branches or just retells.
- ▹Pick one bug this week and debug it with the AI pair turned off end to end — compare how it felt to your usual flow.
- ▹Read Google's SRE book chapter 'Postmortem Culture: Learning from Failure' on blameless postmortems — the same process-over-outcome discipline applies here, just aimed at a new bottleneck.
- ▹If you manage juniors, audit one week of their merged PRs and ask yourself which ones you could confidently say they could reproduce the diagnosis for, unaided.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.