My Engineer Swore the AI Agent Made Him Faster. The Stopwatch Said Otherwise — He Used It Anyway
Day 2 — The Job Behind the Job
Here's the uncomfortable part: I timed a senior engineer using our coding agent, proved it made him slower, put the numbers in front of him — and his usage didn't move, not that week, not the next. If you build or evaluate AI coding agents, this is not a fluke, it's the whole game. People keep hiring a tool for a job they never say out loud, and if you only optimize for the job they claim, you can ship a faster tool and watch adoption sit completely flat.
The experiment: clocked and still losing
One of our senior engineers at PhoenixDX kept insisting the coding agent made him faster. Fine — let's check. Same engineer, same class of task, once with the agent and once without. Without it: done in under twenty minutes. With it: over thirty, once you counted the prompt writing, the diff review, and cleaning up after a hallucinated API call. Not close. Objectively, measurably slower.
We showed him the numbers. He looked at them and agreed — no argument, no denial. Usage did not drop. If you're only measuring the stated functional job — write code faster — that's a contradiction sitting right in front of you. It stops being a contradiction the second you stop assuming that's the job he was actually hiring the tool for.
The JTBD move: stop asking if the tool is good
Jobs To Be Done (JTBD) reframes the question. Instead of "is this tool good" — a product-quality question with no fixed answer, the kind that generates endless survey noise — you ask "what job did you hire it for, in that specific moment." That's causal, not evaluative. You're not rating the tool. You're reconstructing the decision to reach for it.
So we ran it as an actual interview, the way JTBD practice prescribes: sit with the engineer, pull up the exact session from the logs, walk through it chronologically. Not "do you like the agent" but "walk me through this Tuesday night — what was happening right before you opened it."
The answer that didn't fit the spreadsheet
His first answer was the one everyone gives: "it helps me move faster." We didn't stop there. Pushed on what was actually going on that night, and the real answer came out: "it's 11pm, I'm the only senior person online, and I just needed to think out loud with someone who wouldn't judge me for asking something dumb."
There's no spreadsheet column for that. It's not a productivity metric, it's not in any benchmark — and it's the actual reason the tool got opened that night, speed or no speed.
Naming the pattern: functional job vs. emotional/social job
Every tool gets hired for a functional job — the task it visibly does: write code, fix a bug, summarize a doc — and, often at the same time, an emotional or social job: how it makes the person feel, or how it lets them appear to others. Speed is functional. Not feeling alone at 11pm is emotional. Having something to point to if a decision goes wrong — "the agent suggested this approach" — is social cover.
- ▹Functional job: complete the task (write code, refactor, debug) — measurable in time, tokens, correctness.
- ▹Emotional job: reduce anxiety, feel supported, avoid the discomfort of being stuck alone.
- ▹Social job: look competent, diffuse responsibility, avoid looking stuck in front of a manager or teammate.
The functional job is almost always the one people say out loud — even to themselves. Engineering culture rewards reasons that sound rational and efficient. "I use it because it's faster" is a sentence you can say in standup. "I use it because I was lonely and needed reassurance" is not something most engineers will volunteer, or even consciously register, on the first pass.
Why this matters when you're building or evaluating agent tools
If you only optimize for the stated job, your eval suite ends up as task-completion rate, latency, token cost — all real, all functional, and all blind to the reason usage doesn't move when those numbers do move. You can ship a measurably faster agent and watch adoption stay flat, because you never touched the job that was actually driving the behavior.
This is also why chat-style, conversational surfaces keep showing up inside tools whose core value proposition is autocomplete or task automation. It's worth pulling your own usage logs and checking for off-hours or weekend activity that "faster shipping" alone wouldn't explain. If your product decisions and benchmarks only track the functional axis, you're steering with half the instrument panel.
Close: the setup for Day 3
You can't get to the real job by asking "why do you use this" and stopping at the first answer — you'll get the functional one every time, because it's the one that's socially safe and easiest to reach for. The move is one follow-up question: "what would you have done in the five minutes before you opened this tool?" That forces them to describe the actual prior state — stuck, anxious, alone, unsure — instead of the abstract capability they associate with the tool. Day 3 is the full interview script built around that one question.
Extend your knowledge
- ▹Read Competing Against Luck by Clayton Christensen — the foundational JTBD text this framework comes from.
- ▹Look up Bob Moesta's JTBD interview technique (Rewired Group) — the source of the 'walk me through the moment' interview structure used here.
- ▹Pull your own agent tool's usage logs and check for off-hours or late-night spikes — a rough proxy signal for an emotional/social job hiding under a functional one.
- ▹Come back for Day 3: the full interview script for surfacing the real job without accepting the first answer.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.