Why this matters
Your RAG pipeline isn't broken — it's lying to you politely. When it's subtly wrong instead of loudly wrong, the bug is almost never the embedding model. It's the chunker. Nobody checks the chunker, because splitting text into chunks feels like the one part of the stack that's already solved.
The incident: 9 out of 10, then a flat denial
I was demoing a support-docs RAG agent to a client. Nine questions in, nine clean answers, citations lined up like they were supposed to. Then question ten: does the refund policy cover digital goods bought more than 30 days ago? The agent said no — flat, confident, wrong. The docs said otherwise. One sentence, buried in a policy page, gave digital goods a 45-day window instead of 30.
My first instinct was to blame generation — maybe the model just invented a denial. It hadn't. I pulled the actual chunks retrieved for that query, and there it was. The retrieved chunk ended mid-sentence: '...digital goods are subject to an extended return window of' — and stopped. The next chunk picked up with '45 days from the purchase date, provided the item has not been redeemed.' One fact, two vectors, split clean down the middle. The retriever did exactly what it was built to do and pulled the first chunk — decent cosine similarity to the query. The second chunk, the one holding the actual number, never made it into context. On its own, '45 days from the purchase date' doesn't look like it's about anything in particular, so it scored lower and got left behind.
The model wasn't hallucinating. It was faithfully summarizing half a sentence. This is the single most common root cause behind a RAG agent confidently denying something that's sitting right there in the source docs — not a model capability problem, a chunk boundary problem.
Before and after: the actual chunk boundaries
Here's what a fixed-size splitter — 512 tokens, zero sentence awareness — did to this exact text, next to what a boundary-aware splitter does with the same source.
SOURCE TEXT:
"Standard items may be returned within 30 days of purchase for a full
refund. Digital goods are subject to an extended return window of 45
days from the purchase date, provided the item has not been redeemed
or downloaded more than twice."
--- FIXED-SIZE SPLITTER (cuts every N tokens, ignores sentences) ---
Chunk A: "...Standard items may be returned within 30 days of purchase
for a full refund. Digital goods are subject to an extended return
window of"
Chunk B: "45 days from the purchase date, provided the item has not
been redeemed or downloaded more than twice."
--- RECURSIVE / SENTENCE-AWARE SPLITTER (splits on sentence boundary,
with overlap) ---
Chunk A: "Standard items may be returned within 30 days of purchase
for a full refund."
Chunk B: "Digital goods are subject to an extended return window of
45 days from the purchase date, provided the item has not been
redeemed or downloaded more than twice."Same source, same embedding model, same retriever. The only thing that changed is where the cut landed. In the fixed-size version, the fact (45 days) gets severed from its subject (digital goods). In the sentence-aware version, every chunk is a complete, retrievable claim.
Why this is invisible in normal evals
This bug won't show up on the dashboards most teams actually watch. Three reasons why:
- ▹Nothing crashes. No exception, no empty result, no latency spike — the pipeline hands back a perfectly formed, confident, wrong answer.
- ▹Cosine similarity still looks fine on both halves. The chunk ending in '...extended return window of' still scores reasonably against 'digital goods refund policy' because it shares vocabulary with the topic. The embedding isn't wrong — it just can't represent a fact that isn't actually in the text.
- ▹Topic-level relevance metrics — the ones most retrieval evals check first — grade on 'is this chunk about the right subject,' not 'does this chunk contain the specific fact being asked about.' Both halves pass a topic check easily. Only a fact-level check, the kind most teams skip, catches this.
If you're only watching retrieval@k or answer-relevance scores, a split-fact bug and a healthy pipeline look nearly identical. You need an eval that checks whether the specific answer was present, not whether something on-topic got retrieved — which is exactly what's coming tomorrow.
The instinct trap
When teams hit this failure mode, the near-universal first move is to swap the embedding model, bolt on a reranker, or tune top-k. I've watched teams burn weeks there. It's not an unreasonable instinct — chunking feels solved (there's a default `RecursiveCharacterTextSplitter` call in every tutorial, so surely it's fine), while embeddings and rerankers are the parts everyone associates with 'retrieval quality.' So that's where the attention goes by default.
But a better embedding model will embed a severed half-sentence more precisely — it won't un-sever it. Reranking can only reorder chunks that already exist; it can't reassemble a fact that got cut into two non-answers. If the fact isn't intact in any single chunk, no amount of retrieval tuning brings it back. You have to go fix the chunker.
Three chunking failure patterns that cause this
- ▹Mid-sentence splits — a fixed character or token count cuts straight through a sentence, like the refund example above. The subject lands in one chunk, the predicate in another, and neither is retrievable as a complete claim.
- ▹Losing the preceding header or context — a chunk reads '45 days from purchase date' with nothing telling you this is specifically about digital goods, because the splitter cut right after the section heading 'Digital Goods Returns' landed in the previous chunk. The fact is intact, but now it's ambiguous, or reads like it applies to everything.
- ▹Table rows separated from their headers — a pricing or eligibility table gets chunked row by row, so a chunk contains '45 | No | Yes' while the column headers ('Window (days) | Restocking fee | Store credit eligible') sit in a different chunk entirely. The numbers are retrievable. They're also meaningless without their labels.
The fix, and how to audit your own pipeline today
The fix is semantic or recursive chunking that respects sentence and section boundaries, with overlap sized to cover a sentence or two instead of some arbitrary token count. Split on paragraph and sentence boundaries first, fall back to a character limit only when a single sentence blows past your chunk size, and prepend the active section header to every chunk under it — not just the first one. For tables, either keep the header row attached to every data row, or flatten each row into a self-contained sentence ('Digital goods have a 45-day return window, no restocking fee, store credit eligible') before you embed it.
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=400,
chunk_overlap=80, # ~1-2 sentences of overlap
separators=["\n\n", "\n", ". ", " ", ""], # try paragraph, then
# newline, then sentence,
# only fall back to raw
# char split as last resort
)Here's the audit you can run today, before touching a single config value: grep your existing chunk store for chunks that end without terminal punctuation, or that start with a lowercase letter or a mid-word continuation. Those are your split-sentence suspects.
# flag chunks that look like they end mid-sentence
jq -r '.text' chunks.jsonl | grep -nvE '[.?!"'\'')\]]\s*$'
# flag chunks that look like they START mid-sentence
# (starts lowercase, or starts with a continuation word)
jq -r '.text' chunks.jsonl | grep -nE '^[a-z]|^(and|but|or|however|therefore|which|that)\b'Run that against a few hundred chunks from your real corpus. If even 5-10% look like they end mid-sentence, that's your silent-wrong-answer rate, just waiting for the right query to expose it.
Where this goes next
This exact bug class is what a retrieval-level eval is built to catch — not 'is the topic right' but 'is the specific fact the question needs actually sitting in what got retrieved.' Day 22 builds that eval, so next time this happens, you catch it before it happens live, in front of a client, on question ten.
Extend your knowledge
- ▹Read the chunking strategies section of the LangChain text-splitters docs — compare `RecursiveCharacterTextSplitter` defaults against a markdown-header-aware splitter for your own doc format.
- ▹Try LlamaIndex's semantic chunking (`SemanticSplitterNodeParser`, which splits on embedding similarity shifts between sentences) and see how its boundaries compare to your current fixed-size splitter on the same corpus.
- ▹Run the grep audit from this lesson against your actual chunk store today — it takes under ten minutes and tells you whether you have this bug right now.
- ▹Save Day 22 (retrieval evaluation) for when you're ready to turn this audit into an automated, repeatable check instead of a one-off grep.
Discussion
Chat with Chi Cong (AI) about this article. Your conversation is private to you — you can publish a summary for others when you're done.