Day 19: Hybrid Search Isn't a Free Upgrade — It's a Hyperparameter You Have to Tune
Every RAG blog post tells you the same thing: bolt vector search onto BM25, fuse the results, get the best of both worlds for free. Here's what they don't tell you — the fusion step runs on a constant almost nobody touches, and the metric everyone uses to prove the upgrade worked, recall@k, doesn't actually tell you whether your agent got the right answer.
The Slack message
Two weeks after we shipped hybrid search, an on-call engineer pinged me: "hybrid search is worse than what we had before." Our agent — a support and debugging assistant that pulls docs, config references, and API specs to answer engineering questions — started missing things it used to nail: exact error codes, config key names, specific API method names. The stuff plain BM25 used to get right every single time.
What we shipped
The setup was textbook. BM25 handled lexical, exact-match retrieval. A vector index handled semantic similarity. We fused the two with Reciprocal Rank Fusion (RRF), using the default constant — k=60 — that shows up in every library example and every blog post, copied straight from the original paper.
RRF score for a document d =
sum over each retriever r of 1 / (k + rank_r(d))
k = 60 (the default nobody changes)
rank_r(d) = d's rank position in retriever r's result list (1-indexed)RRF's whole selling point is that it's score-free — it only looks at where a document ranked, not how confident either retriever was. That's also exactly where it falls apart. More on that below.
Why it looked fine at first
In offline eval, recall@k went up compared to BM25-only. The right chunk showed up in the top-k more often across our test set. We called it a win, shipped it, and moved on to the next thing on the roadmap.
The gap: recall@k vs. what the agent actually needs
Recall@k answers one narrow question: is the correct chunk somewhere in the top-k? It has nothing to say about where in that list it landed, whether the agent's context window and attention actually used it, or whether the final answer was right. When an agent calls retrieval as a tool and reasons over whatever comes back, rank position and surrounding noise matter just as much as presence.
- ▹Recall@k is a binary, chunk-level, retrieval-only score
- ▹Task success is whether the agent's actual output — the answer, the patch, the action — was correct
- ▹A chunk sitting at position 8 of 10 satisfies recall@10 and still contributes nothing: the model may barely weight it, or a reranker might drop it before it reaches the prompt
- ▹Nobody was tracking task success as a metric until production complaints forced the issue
Diagnosing it
We pulled the failure cases — queries where the agent gave a wrong or incomplete answer after launch. The pattern repeated: queries with exact strings, error codes, config keys, API method names, that BM25 used to rank at position 1 were now landing several spots lower after fusion, or falling out of the agent's usable top-k entirely.
The cause was RRF itself. It only sees rank, never score. BM25 can be dead certain about a match — a near-unique term hit — but RRF flattens that certainty into the same 1/(k+rank) bucket as a vector result that barely cleared the similarity threshold. A rock-solid BM25 hit at rank 1 gets diluted by a so-so vector result also near rank 1, and the fused list loses the one signal that mattered: this wasn't just plausible, it was the answer.
The fix
Two changes, and this time we tested against task success instead of recall@k:
- ▹Pulled the RRF constant down from k=60, which gives top-ranked results more pull and cuts the smoothing that was burying BM25's high-confidence hits
- ▹Swapped pure rank-based RRF for score-aware weighted fusion — normalize BM25 and vector scores, blend them with a tunable weight, so a high-scoring BM25 exact match can actually dominate instead of getting rank-capped
- ▹Reran the eval loop against an LLM-as-judge (or human-graded) task success score on real downstream queries — not recall@k — and tuned the fusion weights against that instead
Task success on the exact-match slice recovered. Recall@k barely budged — which was exactly the point. We'd been shipping changes against a metric that was blind to the regression it was supposed to catch.
The uncomfortable result
Even after tuning, one slice of queries never came back: short, exact-term lookups — error codes, API names — where plain BM25 still beat the tuned hybrid pipeline on task success. Fusion cost us latency, a second index to maintain, and a new hyperparameter to babysit, with no payoff on that slice. Hybrid search wasn't strictly better. It was a trade: a win on semantic, paraphrased queries, a cost on exact-match ones, and the net depended entirely on your query mix and how much tuning effort you were willing to spend.
Takeaway — and the bridge to Day 20
Fusion isn't a one-time architecture decision you ship and forget. It's an ongoing hyperparameter — the RRF constant, the score weights — that has to be tuned against the metric your agent actually lives or dies by, not whichever metric is cheapest to compute. Measure recall@k, declare victory, and move on, and that's exactly how a 'strictly better' hybrid system quietly gets worse in production. Day 20 picks up from here: once fusion is tunable and measured correctly, when is it even worth the complexity — and when does a single well-tuned retriever beat it outright for a given query distribution?
Extend your knowledge
- ▹Read the original Reciprocal Rank Fusion paper (Cormack et al., 2009) and see why k=60 made sense in that context — and why it's not a universal default
- ▹Running an agentic RAG pipeline? Set up an LLM-as-judge or human-graded task success eval next to recall@k before you touch a single fusion weight
- ▹Segment your eval set by query type — exact-match/lexical versus semantic/paraphrased. A single aggregate score will hide exactly the regression described here
- ▹Try weighted score fusion — normalized BM25 plus vector scores with a tunable alpha — as an alternative to rank-only RRF, and benchmark both against single-retriever baselines
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.