Back to blog
Series · Day 20
Data & Retrieval Engineering in 30 Days
View all lessons →
Semantic Search

Day 20 — Semantic Search: A Different Data Model, Not a Smarter Query Box

Every agent you've shipped this year has a retrieval step hiding behind it — RAG, tool memory, context injection, all of it. If you don't actually understand what changes when you swap keyword search for vectors, here's what happens: retrieval quality falls off a cliff in production, and you spend a week debugging the LLM when the bug was three layers downstream, in the index.

The hook: 'cheap laptop' vs 'budget-friendly'

Someone searches your product catalog for "cheap laptop." Your Elasticsearch index has a product tagged "budget-friendly ultrabook." Zero token overlap, zero hits — keyword search just failed a query any human would call trivial. BM25 scores shared-term frequency weighted by rarity; it has no concept that "cheap" and "budget-friendly" are neighbors in meaning. Semantic search catches this because it's never comparing words in the first place — it's comparing positions in a learned meaning space, and "cheap laptop" lands right next to "budget-friendly ultrabook" in that space even though the two phrases don't share a single token.

What actually changes under the hood

This isn't a new ranking formula bolted onto your existing index. It's a different storage primitive and a different matching operation, top to bottom.

  • ▹Keyword search: text → tokens → inverted index (term → list of doc IDs) → match = shared tokens, scored by BM25/TF-IDF.
  • ▹Semantic search: text → embedding model → fixed-length vector (e.g. 768 or 1536 floats) → match = geometric distance (cosine or dot product) between vectors.
  • ▹Inverted index lookup is exact-set intersection — fast and exact. Vector lookup is nearest-neighbor search over continuous space — approximate by necessity once you're past a few thousand vectors.
  • ▹You stop indexing "documents" and start indexing "points in a high-dimensional space." Your old mental model — does this term appear — is gone. It's replaced by: how close are these two points.

Where the vector comes from

An embedding model takes a chunk of text in and hands you back a fixed-length array of numbers. That's the whole contract — no math required to use it. The same model has to encode both your documents at index time and the incoming query at search time, because this whole scheme only works if both land in the same learned space. In 2026 that's usually a hosted API call — OpenAI's embedding models, Voyage, Cohere — or a local model (BGE, E5, Nomic) run through sentence-transformers. Same tradeoff you already make when picking which LLM powers an agent: hosted for simplicity, local for cost and latency control once you're at scale. Pick one model and version per index, and never mix them — a vector from model A means nothing next to a vector from model B, even when the dimensions happen to line up.

Storing and querying vectors: ANN vs a B-tree

A B-tree is built for ordered, exact-or-range lookups, and it's useless in high-dimensional space — "nearest neighbor" isn't an ordering problem, it's a geometry problem. That's the whole reason vector search needs Approximate Nearest Neighbor indexes: HNSW (a navigable graph of vectors you hop across toward the query) and IVF (cluster the vectors into buckets, only search the nearest buckets). Both trade a sliver of recall for a huge jump in speed, because exact nearest-neighbor search over millions of vectors is too slow for anything interactive.

sql
-- pgvector: this is NOT 'just add a column'
CREATE EXTENSION IF NOT EXISTS vector;

ALTER TABLE products ADD COLUMN embedding vector(768);

-- Without this index, every query does a full sequential scan
-- comparing against every row's vector -- fine at 10k rows, dead at 10M.
CREATE INDEX ON products
  USING hnsw (embedding vector_cosine_ops);

SELECT id, name
FROM products
ORDER BY embedding <=> '[0.012, -0.044, ...]'::vector
LIMIT 10;

Here's the trap I see engineers fall into constantly: add a vector column to Postgres, ship it, call it done. Skip the ANN index and you get a query that's correct and linearly scans the whole table — fine in a demo, dead the moment real volume shows up. And HNSW's build time, memory footprint, and the recall/speed knobs (ef_search, m) are now capacity-planning decisions you own, the same way shard counts used to be your problem in Elasticsearch.

The ranking catch: 'close' isn't always 'right'

Semantic similarity measures topical closeness, not relevance — and that gap is exactly where production RAG pipelines quietly lose users' trust. Say a user asks "how do I cancel my subscription," and the nearest vector in your docs is an article titled "How to pause your subscription." Topically adjacent, lexically similar, probably written by the same team — and it does not answer the question that was asked. BM25 would've at least preferred the doc with the literal word "cancel" in it. An agent built on pure vector retrieval hands this wrong-but-close chunk to the LLM as context, the LLM writes a fluent and confident answer, and that's it — that's the real story behind most "why does our support bot hallucinate" tickets. It's not a generation bug. It's a retrieval bug wearing a generation bug's clothes.

Hybrid in practice: BM25 + vectors

Almost no serious production system ships vector-only search. The pattern that actually works is hybrid: run BM25 and vector search in parallel on the same query, then fuse the two ranked lists into one. Reciprocal Rank Fusion (RRF) is the fusion method everyone reaches for, and the reason is simple — it never needs you to normalize two incompatible score scales. BM25 scores and cosine distances live in different ranges and don't average meaningfully. RRF sidesteps that entirely; all it needs is each result's rank position in its own list.

python
def reciprocal_rank_fusion(bm25_ranked_ids, vector_ranked_ids, k=60):
    scores = {}
    for rank, doc_id in enumerate(bm25_ranked_ids):
        scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank + 1)
    for rank, doc_id in enumerate(vector_ranked_ids):
        scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank + 1)
    return sorted(scores, key=scores.get, reverse=True)

Elasticsearch's RRF retriever and OpenSearch's hybrid query — which added RRF as one of its fusion techniques alongside score normalization — implement close to exactly this, and so do Postgres setups pairing pgvector with tsvector/BM25 extensions. For an agent-facing retrieval layer, hybrid search is the cheapest fix for the "close but wrong" failure above: the keyword signal is a guardrail the pure embedding path never had.

Where this fits in the 30-day arc

Today you built the vector store and saw why it's not a drop-in column. Tomorrow builds directly on this: once you can retrieve the right chunks, the next failure mode waiting for you is chunking and context assembly — how you split documents before embedding decides whether retrieval ever had a chance of finding the right unit of text in the first place, and that matters a lot once that retrieved text becomes an LLM's or an agent's actual working context.

Flashcards
Check yourself

Extend your knowledge

  • ▹Read the original HNSW paper (Malkov & Yashunin) for the graph-traversal intuition behind why it's fast without being exact.
  • ▹Try pgvector's hybrid search recipes in the official pgvector GitHub repo — run them against a pure-vector query on the same dataset and watch the fusion effect happen in front of you.
  • ▹Look at how OpenSearch and Elasticsearch implement their native hybrid (BM25 + kNN) query APIs — notice how they treat fusion as a first-class query type, not an afterthought bolted on.
  • ▹If you're building a RAG pipeline, instrument retrieval separately from generation — log retrieved chunks and scores — so a 'wrong answer' bug can be diagnosed as retrieval failure vs generation failure. Same split as the cancel-vs-pause example above.
Test yourself on this lesson →

Discussion

Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.

Ask me anything about “Semantic Search” — trade-offs, decisions, or the story behind it.