Prompt Cache Invalidation
Day 19 — Prompt Caching Is a Contract, Not a Toggle
Prompt caching is the one lever keeping your agent fleet's API bill sane — and it runs on a rule nobody writes down anywhere: the cached prefix has to stay byte-for-byte identical, forever, or the whole thing silently falls apart. Break that rule and nothing alerts you. Output stays perfect, tests stay green, and three hours later someone from finance pings you asking why yesterday's bill tripled.
The hotfix nobody connected to the bill
It starts small, the way these things always do. Someone notices a tool-calling edge case — the model keeps passing a malformed date into one function — and pushes a one-line fix to the shared tool-definition block sitting at the top of the system prompt. No PR ceremony, barely a glance in review. It's a schema tweak, not 'real' logic. It ships in a normal deploy, the edge case disappears, output quality actually improves, and everyone moves on with their day. Four hours later the inference bill is 3x normal. Nobody thinks to check the deploy log first — they check traffic, they check for a runaway retry loop, they wonder if someone kicked off a backfill job overnight. The actual cause, one edited line near the top of a prompt, is the last place anyone looks, because it looked like a clean win.
The mechanism: caching keys on an exact prefix
Prompt caching, as Anthropic implements it (and it's functionally the same across providers), hashes the prompt up to a declared breakpoint and reuses the pre-computed attention state for that prefix on the next call instead of reprocessing it from scratch. That's the whole trick behind making a long, repeated system prompt — tool schemas, few-shot examples, persona instructions — cheap and fast on every call after the first: you pay full price once, then a fraction of that on every cache hit inside the TTL window (5 minutes by default, up to an hour if you turn on extended caching).
- ▹The cache key is the token sequence up to the breakpoint — not a semantic hash, not a diff. Change one character anywhere before that breakpoint and the entire prefix is a miss.
- ▹Position matters more than people expect: anything near the top of the prompt sits inside every downstream segment's prefix. Edit a tool schema or a few-shot example near the top and you invalidate everything after it, even the parts you never touched.
- ▹A cache miss doesn't throw an error. It just quietly falls back to full-price, full-latency processing. No warning, no degraded output — the model usually performs identically or better, which is exactly why the edit shipped in the first place.
That's the trap. Every signal an engineer normally checks right after a deploy — does it still answer correctly, did latency spike on this one request, did the error rate move — comes back clean. The only symptom is cost and average latency drifting upward in aggregate, hours later, across every agent sharing that prefix.
Why it took hours to notice
Cache hit rate isn't on the dashboard anyone's watching in the moment. Teams instrument latency, error rate, and output quality because those are the numbers that page someone at 2am. Cache hit rate is a cost metric — something a finance or platform lead skims in a weekly spend review, not something tied to a specific deploy. The signal is there; most providers hand you cache read/write token counts on every single response. Nobody's alerting on the derivative — the sudden drop in hit rate — only on the absolute dollar figure a week later, by which point it reads as ordinary usage growth.
The multi-agent twist
In a single-agent setup this is an annoying surprise — one team eats one bad week of spend and learns a lesson. In a multi-agent fleet it's a coordination failure dressed up as an infra problem. A shared tool-definition block, a shared retrieval-formatting instruction, a shared persona preamble — these get pulled into one file precisely because every agent in the fleet imports it. That's correct architecture. It's also why the blast radius of a 'minor' edit isn't one agent's cache, it's every agent's cache, all at once, the second the change deploys. The team that made the fix never sees the damage. It shows up on other teams' agents, in other teams' cost dashboards, for a change they didn't make and don't even know happened.
The fix: treat the prefix like a schema
- ▹Version the cacheable prefix explicitly — a tool-schema file or shared prompt block gets its own changelog, not a diff buried inside an unrelated PR.
- ▹Require review for anything edited before the cache breakpoint, same as you'd require for a database migration. The blast radius is the same shape.
- ▹Put cache-hit-rate, and cache read/write token counts, on the same dashboard as latency and error rate, with alerting on a sudden drop — not just a line in a weekly cost review.
- ▹If you control breakpoint placement, push genuinely volatile content — user-specific context, today's date, per-request data — below the breakpoint, and keep the top of the prompt as stable as a public API.
- ▹Before you touch a shared prefix, grep for who else imports it. In a fleet, 'my prompt' is usually 'our prompt.'
Takeaway
Prompt caching isn't a free performance toggle the runtime quietly handles for you. It's a contract between everyone who shares a prefix. The moment two or more agents import the same tool schema or system-prompt block, that block stops being a prompt and becomes shared infrastructure — with all the discipline that implies: versioning, review, monitoring. Know who else is reading your system prompt before you touch the top of it.
Extend your knowledge
- ▹Read Anthropic's prompt caching docs for exact breakpoint rules, TTL options (5-minute default vs. 1-hour extended caching), and how cache read/write tokens are reported in the API response.
- ▹Audit your own fleet: grep for every agent that imports a shared system-prompt or tool-definition file, and write down who they are somewhere visible before the next person edits it.
- ▹Add cache hit rate (or cache read/write token ratio) to your existing observability stack next to latency and error rate — most providers expose it per-response, so it's a dashboard addition, not a new instrumentation project.
- ▹If you're on a provider without explicit cache breakpoints, test it empirically: make a trivial top-of-prompt edit in staging and watch whether your latency or cost per call moves. That tells you where your own effective breakpoint sits.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.