Back to blog

Why polish no longer signals quality in AI-generated work

Oct 5, 2026
Series · Day 4
Using AI Day to Day
View all lessons →
Why polish no longer signals quality in AI-generated work

Day 4: Decoration Is Free Now — Here's How I Decide What to Trust

Every doc, dashboard, and UI an AI touches now comes out looking shipped. Headers in the right place, key terms bolded, a tidy summary up top. That used to mean someone cared enough to finish the thing. It doesn't mean that anymore, and if your trust heuristic still runs on 'looks done,' you will approve the wrong row in a table of 200 and never know it.

Two documents, one tell

Last month I had two design docs open side by side. Same headers, same bolded terms, same tidy bullet summary at the top, same 'Key Decisions' callout box. One came from an engineer who'd spent two days thinking through a migration. The other was Claude's first pass at a half-formed Slack thread. Looking at them, I could not tell which was which. I had to read both start to finish before I found it: one of them contradicted itself on page two. That's the moment 'looks finished' stopped meaning anything to me.

The mechanism: polish used to cost something

Formatting used to be a proxy for effort. Someone who structured a doc with headers, bolded the key terms, built a comparison table, had spent time thinking about how to communicate the thing — and that correlated, imperfectly but reliably, with having thought about the content itself. An agent now produces that same structure in the time it takes to generate a paragraph. The structure is free. And a signal that costs nothing to fake stops being a signal. So I flipped the heuristic: polish is now a reason to slow down and check, not a reason to relax.

This isn't a claim that AI output can't be trusted. It's narrower than that: the one cheap proxy I used to lean on — does this look like someone finished it — got arbitraged away. I needed proxies that are expensive to fake. Three of them, below, are what I actually use day to day.

Choice 1: I print what matters and read it start to finish

When something actually matters — an architecture decision, an incident writeup, a PRD I'll be held to — I print it, or at minimum read it in a single pane, top to bottom. No tabs, no AI summary, no keyword search. In 2026 this feels almost willfully inefficient, since I could ask the same assistant that wrote the doc to summarize it for me. That's exactly the problem: summarizing it with the tool that generated it just launders the polish problem one more time.

What linear reading buys me is spatial memory of the argument. I remember 'the risky assumption was in the third section, right after the rollback plan' — not because I tagged it, but because I physically passed it in that order. Search finds a string. Reading in order lets you notice that the string you'd search for later contradicts something two pages back — exactly the kind of error an AI-assisted doc produces, because the model is coherent paragraph by paragraph and only globally coherent by luck.

Choice 2: plain tables, every row visible, every column comparable

The second filter is for anything with more than a handful of data points: a cost comparison, a list of agent runs, a dependency audit, eval results across models. I want one flat table. Every row visible. Nothing collapsed, nothing behind a hover, nothing behind someone else's decision about which rows are 'important enough' to show by default.

Generated dashboards love to hide exactly this. Ask an agent for eval results across five models and ten prompts, and it will often hand you a card layout, a 'top performers' summary, maybe a chart with the worst run smoothed out of frame because it's an outlier. Someone — or something — made a judgment call about what you need to see, and you didn't get a vote. A plain table with 50 rows and 10 columns is unglamorous, and it's the only format where the one wrong row — the eval that regressed, the agent run that silently retried four times, the dependency with a CVE — can't hide behind the furniture. If you can't see all of it at once, you can't audit it. And an AI-generated summary of a table is just a second model's opinion stacked on top of the first model's output, with nothing underneath you can check.

Choice 3: fixed components, not a fresh screen every time

The third filter is specific to UI, and it's the one that bit me hardest. Tools like v0, Claude's artifacts, and Cursor's agent mode produce a genuinely decent-looking internal screen in one prompt — a new admin page, a new eval viewer, a new agent-fleet monitor. No single screen looks bad. The problem shows up at screen twelve: a different card shadow, a different date format, a different definition of 'status: failed' than screen three, because each one got generated independently with no shared vocabulary between them.

A fixed, small component set — the same badge, the same table, the same empty-state pattern, reused everywhere — is boring on purpose. Boring is reviewable: a teammate can look at a diff and tell whether this screen uses the component correctly, the same way they'd review any other code. A one-off generated screen can only be reviewed as itself, from scratch, every single time, because there's no shared pattern to check it against. When I ask an agent to build an internal tool now, I point it at the existing component library and tell it explicitly not to invent new patterns. The constraint is the point, not a limitation I'm working around.

Where fancy is still the right call

None of this means plaintext everything. A customer-facing landing page, a stakeholder demo, a one-off visualization you'll look at exactly once to answer one question — polish is correct there, and AI is genuinely fast at producing it. The filter isn't form versus function. It's how many times you'll touch the thing. Something you'll read once and discard can afford to be generated fresh and pretty. Something you'll read, compare, debug, or maintain needs a format stable enough that your second look is actually comparable to your first.

The test to run on the next thing that crosses your desk

  • ▹Can I see all of it at once — no tabs, no collapsed sections, no 'load more'?
  • ▹Can I compare it — row against row, screen against screen, run against run — in a format that doesn't change shape each time?
  • ▹Can I fix it myself — edit the table, swap the component, correct the doc — without regenerating the whole thing and hoping the polish survives?

If any answer is no, the formatting is doing work that should be yours. Flatten it back down before you trust it.

Closing the loop on this series

Days 1 through 3 of this series were all accelerator: using AI to generate more docs, more code, more UI, faster than you could alone. Day 4 is the brake pedal. It's not about using AI less. It's about knowing which outputs are allowed to look finished without earning it, and keeping the rest in a format that can't lie to you by looking good.

Flashcards
Check yourself

Extend your knowledge

  • ▹Audit one dashboard or internal tool you rely on weekly: count how many rows or states are hidden behind a card, a hover, or a 'show more' — then ask whether you'd have caught last week's anomaly if you only saw what's visible by default.
  • ▹Pick your team's most-reused internal UI (admin panel, eval viewer, agent monitor) and check whether every screen uses the same badge, table, and status components — or whether each one was generated fresh and quietly drifted.
  • ▹Next time an agent hands you a design doc or incident writeup, read it fully in one linear pass before asking any tool to summarize it — compare what you noticed against what the summary would have told you.
  • ▹Revisit Simon Willison's writing on reviewing AI-generated code for a related take: the review process, not the output's appearance, is what has to scale with AI-generated volume.
Test yourself on this lesson →

Discussion

Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.

Ask me anything about “Why polish no longer signals quality in AI-generated work” — trade-offs, decisions, or the story behind it.