How to spot unfalsifiable claims in your writing
Day 1: the sentence that can't be wrong
Somewhere in your last ten PR descriptions is a sentence that could never turn out to be wrong. Not because it's true — because it's built that way. You write these things every day, usually with an LLM open in another tab, and if you can't tell a checkable claim from a decorative one, you have no way of knowing whether the AI just helped you communicate, or just helped you sound confident.
A sentence that sounds finished
Here's a real-shaped PR sentence: "This PR improves reliability and streamlines the deployment process." Read it again. It sounds done. Nothing in it is technically wrong. Now do the exercise: what would have to be true in the world for this sentence to be false? Deploys got slower — is it false? No, 'streamlines' is vague enough to survive that. The service crashed the next day — is it false? Still no, because 'improves reliability' never said how much, or measured against what. There's no observation that could contradict it. That's the tell.
Name the pattern
Polish is decorative — it survives any outcome. 'Drives meaningful impact,' 'seamlessly integrates,' 'robust and scalable': all three are compatible with success, failure, and everything in between, which is exactly why they say nothing. Boring, specific terms are load-bearing — 'p95 latency,' 'retries three times,' 'fails at step 3' — because someone can check them, and someone can be wrong about them. A sentence only earns its place if it could turn out to be false. Not a style preference. The entire difference between information and noise.
Two tests, literally
- ▹Test 1 — Delete it. Cut the sentence from the PR description. Did the reader lose anything they actually needed? If not, it was filler.
- ▹Test 2 — Could it be false? Try to picture the concrete case that would make the sentence wrong. Can't picture one? It isn't a claim, it's atmosphere.
Side-by-side: AI-flavored vs. checkable
- ▹"This change improves reliability and streamlines deployment." → "Cuts deploy time from 12 min to 4 min by parallelizing the build step; retries the flaky step-3 integration test up to 3 times before failing the pipeline." (swap: 'improves/streamlines' → actual numbers and a named failure mode)
- ▹"The new caching layer significantly boosts performance." → "p95 latency for /search dropped from 340ms to 90ms after adding a 5-minute TTL cache in front of the ranking service." (swap: 'significantly boosts' → a metric, a baseline, and a mechanism)
- ▹"Our agent architecture is robust and scales seamlessly." → "The orchestrator retries a failed subagent call twice with exponential backoff, then falls back to a single-agent pass; tested up to 40 concurrent runs before queue latency crossed 2s." (swap: 'robust/seamlessly' → the actual fallback behavior and its breaking point)
- ▹"This refactor drives meaningful improvements to code quality." → "Cyclomatic complexity of the billing module dropped from 28 to 11; test coverage went from 54% to 81%." (swap: 'meaningful' → two numbers anyone can rerun and verify)
Why models drift this way
LLMs are trained on oceans of text optimized to sound finished — marketing copy, exec summaries, polished changelogs — then fine-tuned with RLHF, which tends to reward confident, agreeable output over risky, checkable claims. A concrete claim can get contradicted by a human rater or a later fact. A vague one almost never does. 'Drives meaningful impact' is a safe local optimum, and it's the default register you'll get out of Claude, Copilot, or ChatGPT unless you push back on it directly.
Why we copy it back
It's also socially safer. Nobody can push back on 'this improves reliability' the way they can push back on 'p95 dropped from 340ms to 90ms' — the second one invites 'did you check that on a Friday deploy?' The first just sits there, unfalsifiable and unchallengeable. And once you've read enough model output in your daily PRs, standups, and review comments, that register starts to feel like the normal way to sound competent. It leaks into your own writing even when no model touched the sentence.
Today's practice
Pull up your last three PR descriptions or standup updates. Run the 'could this be false?' test on every sentence. Count how many fail — how many are compatible with any outcome at all. That count is your baseline. You're not chasing zero. You're trying to notice the pattern before you write the next one.
Why this is lesson 1
Everything later in this series — prompting an agent well, reviewing an agent's PR description, delegating a task and trusting the self-report that comes back — rests on this one skill first: telling a checkable claim from a decorative one. When an agent tells you 'tests pass and the fix is solid,' you need to already know that sentence fails both tests before you can ask the question that actually matters: which tests, and what did they check?
Extend your knowledge
- ▹Reread your last postmortem's 'root cause' section and apply the delete test to every sentence in it.
- ▹Read George Orwell's 'Politics and the English Language' — the same rot in political prose, decades before LLMs.
- ▹Next time Claude or Copilot drafts a PR description for you, run the falsifiability test on it before you post it.
- ▹Ask your AI assistant directly: 'rewrite this with only checkable claims' and compare the two versions side by side.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.