Function calling as grammar constraint
Day 17: Function Calling Is a Grammar, Not a Brain Upgrade
Here's the trap I see almost every team fall into: you bolt a JSON schema onto your tool definitions, the model stops spewing garbage, and somewhere in there you quietly start believing it got smarter. It didn't. Schema validation catches malformed output before it hits your code. It has zero opinion on whether a perfectly well-formed call was the right call. When your agent picks the wrong tool next week, you'll go looking in the wrong place if you don't understand that distinction.
The two-day bug that wasn't a reasoning bug
A support-ops team I know had an agent botching roughly one in six refund requests — cancelling orders it should've refunded. The obvious suspect: the model doesn't get the domain. So two days went into few-shot examples, system prompt rewrites, the whole 'it's confused about our policy' theory.
Then someone actually read the logs. The tool's `action` field was a bare string. The model was happily emitting `"refund"`, `"cancel_and_refund"`, and `"void_transaction"` — three different strings, all valid JSON, all sailing through validation, all accepted by the handler. Problem: the handler only knew how to route two of them. `"cancel_and_refund"` fell through to the cancellation path. The moment they tightened the schema to a 3-value enum with real descriptions, two-thirds of the 'reasoning failures' just disappeared. Nobody touched the prompt. One case was left standing — the model picked `"cancel"`, a perfectly legal enum value, when `"refund"` was correct. That one was real. It had been hiding inside a stack of format bugs the whole time.
Grammar vs judgment: draw the line
Function calling constrains syntax — the shape of a valid call: which fields exist, what type each one is, which values are legal. That's a grammar. It says nothing about semantics: whether the call reflects the right decision for the situation in front of the model. I keep seeing teams conflate 'the model emitted a valid, well-typed tool call' with 'the model reasoned correctly.' Those are two unrelated properties. Schema design buys you exactly one of them.
// BEFORE — syntactically fine, semantically ambiguous
{
"name": "adjust_order",
"parameters": {
"type": "object",
"properties": {
"action": { "type": "string" },
"amount": { "type": "number" }
},
"required": ["action"]
}
}
// model can legally emit "refund", "cancel_and_refund", "void_transaction" —
// three strings, three different downstream behaviors, all pass validation
// AFTER — same decision space, grammar now pins it down
{
"name": "adjust_order",
"parameters": {
"type": "object",
"properties": {
"action": {
"type": "string",
"enum": ["refund", "cancel", "void"],
"description": "refund: money returned, order stays open. cancel: order terminated, no money moves. void: same-day reversal before settlement."
},
"amount": { "type": "number", "description": "required when action=refund" }
},
"required": ["action"],
"allOf": [
{ "if": { "properties": { "action": { "const": "refund" } } }, "then": { "required": ["amount"] } }
]
}
}
// the model can STILL choose the wrong enum value — that's the judgment
// bug the schema was never going to fixHow schema validation actually works
The mechanism differs by vendor; the ceiling is identical. OpenAI's structured outputs mode and most open-source grammar-constrained decoders (Outlines, llama.cpp grammars, Guidance) go furthest: they constrain the token distribution at every decode step against the JSON schema. The instant the model would emit a token that breaks the schema — wrong type, a field that doesn't exist, an enum value that isn't on the list, invalid JSON syntax — that token's probability gets zeroed out before sampling. That's a hard constraint enforced by the decoder, not a soft habit the model picked up in training. Anthropic's tool use takes a different route: no token-level grammar lock, just a model trained hard on well-formed calls plus validation on the output afterward — which is also why a malformed Claude tool call occasionally slips through in a spot where OpenAI's strict mode would have made it physically unsamplable. Either way, the ceiling doesn't move: schema enforcement governs the shape of the output, and has no opinion on which of the still-legal tokens or values was the correct decision.
- ▹Schema validation reliably fixes: malformed JSON syntax, wrong field types (a string where a number was expected), hallucinated or unknown parameter names, missing required fields, invalid enum values
- ▹Schema validation does nothing for: the wrong tool chosen among several valid options, the right tool with semantically wrong (but type-valid) arguments, stale or incorrect understanding of system state, policy violations that are perfectly well-typed
The test: run it on your own bug backlog
Take every tool-use failure from the last month and ask one question of each: would a stricter schema — tighter types, an enum instead of a free string, a required field, a discriminated union keyed off an action type — have caught this before it ever reached your code? Or did the model emit a perfectly valid, well-typed call that was still the wrong call? Sort the backlog into those two piles for a week before you fix anything. Early in a project the pile skews hard toward schema bugs, because schemas start loose. As schemas mature, the ratio flips — what's left is almost entirely judgment failures, and no amount of further schema tightening touches them. You'd be polishing a layer that already stopped being the bottleneck.
Day 17 in the multi-agent arc
In a multi-agent system, tool schemas stop being just 'the API into your backend' — they become the contract between agents. One agent's output is the next agent's tool-call input. Typed handoffs, which I'll get into in a later lesson in this series, lean entirely on this grammar-constraint property: agent B can trust that agent A's handoff payload is structurally well-formed, every field present, every type correct. But a contract only enforces shape. If agent A hands agent B a syntactically perfect, strategically wrong plan, agent B inherits a judgment failure dressed up as a valid message. Nothing in your validation logs flags it. It shows up downstream looking like a brand-new bug in agent B, when the actual defect shipped upstream.
The practical takeaway
When an agent misbehaves, resist the reflex to 'improve the schema' as the first move. Diagnose first: is this a shape problem or a choice problem? Tighter types, enums, required fields — those fix shape problems fast, almost for free, and the fix sticks. Choice problems need evals, better context, better prompting, sometimes a different model entirely. Schema changes won't move that needle, and every hour spent tightening a schema against a judgment bug is an hour not spent on the bug that's actually sitting there.
Extend your knowledge
- ▹Read Anthropic's tool use documentation on JSON schema definitions and how tool_choice interacts with the model's output — note that it's enforced through training and output validation, not token-level grammar constraints.
- ▹Read OpenAI's structured outputs documentation for the specifics of how its constrained decoding is implemented against a JSON schema.
- ▹Pick one production tool schema this week and audit every free-string field: which ones are secretly enums in disguise, and what would a discriminated union catch that a flat schema doesn't?
- ▹Preview forward: the upcoming lesson on typed handoffs between agents builds directly on this — the same grammar/judgment split applies to inter-agent contracts, just with another agent reading the schema instead of your backend.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.