Code Knowledge Base Staleness
Day 20 — Your Code KB Isn't a Search Problem, It's a Staleness Problem
If your agent's retrieval comes back textbook-correct and still wrong, stop blaming the embedding model. Blame the clock. Most 'knowledge base for code' failures I've run into aren't retrieval failures at all — they're staleness failures. Nobody wired re-indexing to merge events, so the KB sits there confidently describing a repo that already moved on.
The incident
On a PhoenixDX repo, an agent suggested a code block calling a helper function that had been deleted three commits earlier. This wasn't a fluke retrieval — the chunk came back with a high similarity score, the right file, the right surrounding context, everything you'd want from a textbook RAG hit. The PR survived two rounds of review before someone actually ran the function locally, hit an ImportError, and went digging.
Why this isn't a chunking or embedding problem
Run the same query against the same index today and you'd get the identical result, because retrieval did exactly what it was built to do: find the chunk most semantically similar to the query. That chunk was accurate. Just not for the repo that exists anymore. No amount of better chunking, a bigger embedding model, or reranking fixes this, because the broken part isn't the similarity computation. The index describes the repo as of commit X. HEAD has since moved to X+3. Nothing told the index that.
The real culprit
Most RAG pipelines for code get built on top of RAG pipelines for documents — PDFs, wikis, support tickets. Those systems treat a chunk as a static fact: true when written, still true now. Code doesn't work that way. Code is a living thing with its own clock — the merge-event clock. Every merge to main is a tick. A chunk indexed 40 merges ago isn't 'a little old' — it may describe a function, a signature, a dependency that is simply gone.
- ▹Retrieval failure: the wrong chunk comes back for the query. Fix with better embeddings, chunking, or reranking.
- ▹Staleness failure: the right chunk comes back, but the code it describes no longer exists or no longer behaves that way. Fix with re-indexing and recency signals, not retrieval tuning.
- ▹The incident above passes every retrieval metric you'd check — precision@k, MRR — and still ships a broken PR, because none of those metrics measure time.
The missing piece
Most code-KB setups have no TTL on indexed chunks, no re-index trigger tied to git events, and no signal saying 'this result is six merges old, trust it less than one from HEAD.' The index gets built once, maybe refreshed on a nightly cron, then trusted uniformly ever after — a chunk from the very first index load carries the same authority as one indexed five minutes ago. Once an agent is drafting PRs on its own, that uniform trust is the actual vulnerability. Not the embedding model.
The fix PhoenixDX shipped
Two changes, both deliberately boring. First, the indexer stopped being cron-driven and became event-driven — a webhook on merge-to-main kicks off re-indexing of the changed files within minutes, not hours. Second, every chunk now carries a commit_distance field, merges since that chunk's content last changed, and retrieval scoring down-weights or flags chunks past a threshold. The agent doesn't just get 'here's a matching chunk.' It gets 'here's a matching chunk, and it's 11 merges stale, verify before you use it.'
# merge webhook handler
def on_merge_to_main(payload):
changed_files = payload["commits"]["files_changed"]
reindex(changed_files, commit_sha=payload["after"])
bump_commit_distance_for_all_other_chunks()
# retrieval-time trust signal
def score_chunk(chunk, query_similarity, decay_rate=0.08):
staleness_penalty = 1 - min(chunk.commit_distance * decay_rate, 0.9)
confidence = query_similarity * staleness_penalty
return {
"text": chunk.text,
"similarity": query_similarity,
"commit_distance": chunk.commit_distance,
"confidence": confidence,
"flag": chunk.commit_distance > 15, # surfaced to agent/reviewer
}Before / after
Same query, same embedding model, same top-ranked chunk. The only thing that changed is what's attached to it. Before: the agent gets a bare text chunk and a similarity score, treats it as ground truth, and writes a call to a function that no longer exists. After: the agent gets that same chunk plus commit_distance: 11 and flag: true, and either re-fetches the live file before using it or surfaces the uncertainty to the human reviewer instead of quietly shipping it. Retrieval didn't get smarter. The trust model around it did.
The reframe
A knowledge base for code is a staleness-management problem first and a search problem second. If you're debugging a code-KB failure by swapping embedding models or tuning chunk size, you're working the wrong layer — check whether re-indexing is even wired to git events before you touch retrieval at all. This matters more by the week, as agents move from 'suggest code' to 'open PRs on their own,' because a stale-but-confident retrieval stops being a suggestion a human skims past. It becomes an action an agent takes. Tomorrow we carry this same clock problem into multi-agent systems, where the question isn't just 'is my KB stale' but 'are two agents right now working off different snapshots of the same repo.'
Extend your knowledge
- ▹Audit your current code-KB: is re-indexing cron-based or event-based, and how wide can the gap get between a merge and an index refresh?
- ▹Running a vector DB — pgvector, Pinecone, Weaviate? Check whether it supports per-chunk metadata filtering. That's what lets you add commit_distance scoring without a full re-architecture.
- ▹Look at how your git host exposes changed files on merge — GitHub's push/pull_request events, GitLab's Merge Request Hook. That's the trigger you wire the re-indexer to.
- ▹Read up on cache invalidation from distributed systems. Staleness-vs-freshness tradeoffs in code KBs are the same problem, just wearing different vocabulary.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.