Validating your LLM judge
Why this matters
If an LLM-as-judge validator sits on your critical path, it's production code. And unvalidated production code fails silently, in exactly the direction that hurts you most: it keeps saying yes.
The incident: a timeout that learned to say yes
A customer flagged a batch of malformed outputs — fields missing, one response in the wrong language entirely. We pulled the logs expecting a model regression. Instead we found three weeks of outputs that had been silently approved by our own judge. Usage had grown, the judge model's p95 latency had crept past our timeout, and every timed-out call fell through to this line.
async def judge_output(output, criteria):
try:
verdict = await llm_judge.evaluate(output, criteria, timeout=5)
except TimeoutError:
# written in a hurry, to stop a slow judge from blocking the pipeline
return {"pass": True, "reason": "judge_timeout_fallback"}
return verdictWhoever wrote that fallback wasn't being careless — they were unblocking a pipeline mid-incident and meant to come back to it. Nobody did, because nothing in the dashboard distinguished "the judge approved this" from "the judge never ran." Both said pass.
Name the pattern
This is what happens when a team builds an LLM-as-judge validator and files it mentally under plumbing instead of under feature. Plumbing gets wired up and left alone. A feature gets a spec, a test suite, a version number, and an owner who gets paged when it misbehaves. Most judge prompts get none of that — they're a function call bolted onto the end of a pipeline, trusted by default because they sound authoritative. An LLM confidently returning {"pass": true} reads the same whether it evaluated your output carefully or whether a one-line except block made the decision for it.
Why it's invisible
A dashboard built around pass/fail counts has no third state for "didn't actually check." Fail-open failure modes are specifically the ones that don't show up as failures — that's the whole shape of the problem. The judge call errors, times out, hits malformed JSON it can't parse, or gets rate-limited, and in every one of those cases the pipeline has to decide what "unknown" means. If the default is pass, unknown and good become indistinguishable in every metric you're already watching.
Both branches land on the same green checkmark. The only way to tell them apart is to log the branch itself as a distinct event — which almost nobody does on day one, because day one is about proving the judge works at all, not about proving it fails safely.
The fix: treat the validator like critical-path code
- ▹Unit tests with known-bad fixtures — inputs you've hand-labeled "should fail," run against the judge on every deploy, so a prompt change or model swap that softens the judge gets caught before production does.
- ▹Fail closed, not open — on timeout, error, or an unparseable verdict, the default is reject (or route to human review), never pass. A slower pipeline beats a silently broken one.
- ▹Separate alerting for judge-call failures vs. content failures — "the judge said no" and "the judge couldn't run" are different incidents with different owners; merging them into one pass-rate metric is how this bug hid for three weeks.
- ▹Version the judge prompt like you version a model — tag it, diff it in review, re-run your fixture suite against every change. A one-line prompt edit can quietly shift the approval bar.
Widening the lesson
The deeper mistake wasn't the timeout line — it was believing that "we have a validator" was a finished sentence. Validation isn't a single gate bolted onto the end of a pipeline; it's a component with its own correctness bar, its own failure modes, and its own blast radius when it's wrong. Right now you're probably only checking the final output. Day 19 pushes this further: in a multi-agent pipeline, bad output from agent 2 that slips past an unvalidated check becomes bad input to agent 3, and now you're debugging a compounding error with no clear origin. Same fix, applied earlier — validate the handoffs between agents, not just the finish line.
This week's checklist
- ▹Find every except/timeout/fallback branch in your judge code and check what it returns — if it's pass, that's your bug waiting to happen.
- ▹Write 5–10 known-bad fixtures and run them through your judge today, not as a one-off, but wired into CI.
- ▹Add a metric or log line that fires specifically on "judge call failed," separate from "judge said fail."
- ▹Put a version tag or commit hash on your judge prompt and diff it the next time someone edits it.
Extend your knowledge
- ▹Read Anthropic's 'Building Effective Agents' guide — its evaluator-optimizer pattern is this same problem from the other side: putting an LLM check inside the loop, and what it takes to treat that check as a real component instead of an afterthought.
- ▹Try a prompt/eval testing tool like promptfoo or Braintrust to build a fixture-based regression suite for your judge prompt instead of hand-rolling one.
- ▹Audit one validator you already have in production this week: find its error-handling branch and check whether it fails open.
- ▹Look into OpenTelemetry-style structured logging to tag judge-call outcomes (success/timeout/parse-error) separately from the verdict itself.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.