Tool Schema Versioning
Day 18 — Your Tool Schema Is a Contract, Not a Comment
You rewrote one line of a tool description. Tests green, types green, reviewer approved it in ninety seconds because it was "just a docstring." Three days later your ticket queue is full of garbage and nobody traces it back to that PR. That's the bug this lesson is about — the one change your entire pipeline is built to miss.
The incident: a 'copy tweak' that wasn't
Take a tool called create_refund. It has a param reason: string, described as "Why the refund is being issued." Someone decides that's vague and rewrites it to "Internal note for the support ticket, max one sentence." No type change. No field added, none removed. The schema still validates, the PR reads like a docstring edit, it gets a rubber-stamp approve. Next day, the agent starts stuffing full customer complaint transcripts into reason instead of the short category label your ticketing system expects downstream. Nothing crashes — it's still a string — but the queue fills with noise, and it takes a week before anyone connects it to yesterday's harmless-looking wording change.
Why none of your safeguards catch this
Every layer of defense you already trust was built to catch a different species of bug, and this one walks straight under all of them.
- ▹Unit tests pass, because the tool's implementation never changed — only the schema text the model reads changed.
- ▹Type checkers pass, because "string" is still "string." JSON Schema validation has zero concept of semantic drift.
- ▹Code review waves it through, since it's English prose, and prose edits read as harmless by default. Reviewers scrutinize logic diffs; they skim adjectives.
- ▹The only thing that actually moved is what the model infers from the words — and that inference step stays invisible until it produces an argument shape your code never expected.
Reframe: the schema is the interface, not the documentation
In an ordinary codebase, a docstring is advisory — the compiler enforces the real contract, and a confused human can always go read the source. In agentic systems, the model is the caller, and the JSON schema plus its description fields are the only thing it ever sees. No access to your source, your business logic, your intent — just the contract. So anything that nudges the model's reading of that contract — a reworded sentence, a reshuffled enum, a loosened type — is a breaking interface change, full stop, even when nothing a linter checks has moved an inch.
The parallel backend teams already know
Every team that's shipped a public REST or gRPC API already has the muscle memory for this exact problem. It just hasn't occurred to them yet to point it at tool schemas.
- ▹Semantic versioning: teams that follow it bump a major version for anything that could break a caller's assumptions. Most tool schemas carry no version field at all.
- ▹Deprecation windows: old fields stay live next to new ones while callers migrate over. Tool schemas get edited in place — the old behavior is just gone the second the PR merges.
- ▹Changelogs: API consumers get a diff to read before they upgrade. Agents get nothing — the model discovers the change by making the wrong call, in production, with real data.
- ▹Contract tests: API teams snapshot request/response shapes against the spec. Almost nobody snapshot-tests model behavior against a schema edit.
What breaks, ranked by how silently it fails
Not every schema edit is equally dangerous. Rank them by how loudly they fail — the quiet ones are the ones that actually get you.
- ▹Renaming a param is the loudest failure there is. The old name vanishes; calls using it error out or get rejected on validation immediately. You find out within the hour.
- ▹Narrowing an enum is medium-loud. Calls that hit a now-removed value fail validation, but only for that slice of traffic — so it can hide for days in a low-volume path nobody's watching closely.
- ▹Changing the default behavior implied by prose is quiet. The schema shape hasn't moved, but the model now assumes a different fallback — "omit this field to use the account default" silently turns into "omit this field to use zero" — and your code just processes whatever it's handed.
- ▹Rewording a description that anchors a judgment call is the quietest and most dangerous of the four. Soften "use the strict matching mode" to "use matching mode when appropriate" and nothing structural changes — but WHEN the model bothers to set that field at all shifts. No error, no validation failure, just a statistically different pattern of calls that takes forever to notice.
The fix: treat the schema like you'd treat any service contract
- ▹Pin a schema version per agent/prompt. Even a bare integer field — tool_schema_version — gives you a way to trace a behavior regression back to the exact edit that caused it.
- ▹Snapshot-test model behavior against schema diffs. Run a fixed set of representative prompts through the model on the old schema and the new one, then diff the resulting tool calls — not the schema JSON. That's the test a unit test can't give you, because what's actually under test is the model's inference, not your code.
- ▹Add a contract review pass, separate from code review. Someone whose job is to ask "how could a model misread this sentence," not "does this compile." Treat it the way you'd treat sign-off on an API spec change, not a typo fix.
- ▹Keep a changelog for tool schemas the same way you'd keep one for a public API, even an internal-only one. The next engineer who touches reason: string six months from now needs to know it's load-bearing.
Why this matters for the series
Function calling handed LLMs a type system sitting right at the boundary where natural language meets deterministic code. That boundary deserves exactly the discipline you'd demand at any other service boundary — versioning, review, tests — because the model will never throw an exception to tell you it misunderstood. It'll just call the tool, confidently, with the wrong shape, and your system will do precisely what you told it to do with data you never meant to send it.
Extend your knowledge
- ▹Pull up a live tool schema in your codebase and read every description field as if you were the model with zero other context — mark every spot where a judgment call is implied instead of specified.
- ▹Add a version field to one high-traffic tool schema and write a single snapshot test that runs 5-10 representative prompts through it, capturing the resulting tool calls as a baseline to diff against future edits.
- ▹Look at how your team reviews OpenAPI or protobuf changes today, and ask which of those review steps are completely absent from how tool schema PRs get merged.
- ▹Read Anthropic's and OpenAI's tool-use documentation side by side and compare how explicit each is about treating descriptions as part of the contract versus cosmetic text.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.