The prototype indexed forty curated PDFs and answered beautifully. The production system indexes forty thousand documents — support tickets, contracts with three revisions each, spreadsheets exported to text, a decade of wiki pages nobody has reviewed — and the answers fall apart. The instinct is to blame the model and reach for a bigger one.
That is almost always the wrong diagnosis. Retrieve the chunks for a failing question and read them. Nine times out of ten the passage containing the answer is not in the set. The generator did not hallucinate out of malice; it was asked a question and handed context that did not contain the answer, and it did what such models do — produced something fluent and plausible from what it had.
RAG failures are retrieval failures wearing a generation costume. Once you accept that, the work becomes tractable, because retrieval is a measurable engineering problem with well-understood levers.
Chunking destroys more answers than any other stage
The default in every tutorial is to split on a fixed character count with some overlap. It is the wrong default for real corpora, and the damage is silent.
Split a 900-token contract clause at 512 characters and the obligation ends up in one chunk while the party it binds ends up in another. Neither chunk answers "what is the supplier required to do", and the one that scores best is actively misleading. Split a table and every row loses its header, so a chunk reading | 4.2 | 18 | 2031 | is retrievable and meaningless.
What works is structure-aware splitting: respect the document's own boundaries — headings, sections, list groups, table units — and only fall back to size-based splitting inside a section that genuinely exceeds your limit. For tables, serialise each row with its headers repeated so the row is self-describing. For long sections, keep a parent-document pointer: embed the small chunk for retrieval precision, then expand to the surrounding section before sending to the model, so you get precise matching with sufficient context.
Metadata matters more than chunk size
Teams spend weeks tuning chunk size from 512 to 768 and gain very little. The same weeks spent on metadata change the system's behaviour entirely. Every chunk should carry, at minimum:
- Source identity and a stable deep link — without this you cannot cite, and an uncitable answer is unusable in any serious setting.
- Section path — "Master Services Agreement › Schedule 2 › Termination" is signal for both retrieval and the reader.
- Effective date and version — the difference between the current policy and the one it replaced. Both are in the index. Only one is the answer.
- Access-control identifiers — the groups or roles permitted to see this content, resolvable at query time.
- Document type — a support ticket and a signed contract deserve different trust weights.
Staleness is the failure that erodes trust fastest. A system that confidently quotes a superseded policy is worse than no system, because the answer is well-formed and wrong. Version and effective-date filters at retrieval time are not a refinement; they are load-bearing.
Pure vector search fails on exactly what users ask about
Dense embeddings capture semantic similarity. That is their strength and it is precisely why they fail on identifiers. Error code ERR_5521, part number BX-40-221, the surname of a rarely mentioned counterparty, a specific SKU — these are tokens where you need exact lexical matching, and an embedding will cheerfully return ERR_5522 as a near neighbour because the two strings are semantically almost identical and materially different.
Real user queries are full of these. Support questions are mostly error codes; legal questions are mostly proper nouns. A dense-only system is structurally bad at its most common query type.
The fix is hybrid retrieval — run BM25 and dense search in parallel and fuse the result lists. Reciprocal rank fusion is the right default because it combines by rank rather than by score, so you never have to normalise two incomparable scoring scales:
from collections import defaultdict
def reciprocal_rank_fusion(result_lists, k=60, limit=50):
"""Fuse ranked lists by rank position. k=60 is the standard damping
constant; it keeps any single list from dominating the head."""
scores = defaultdict(float)
chunks = {}
for results in result_lists:
for rank, chunk in enumerate(results):
scores[chunk.id] += 1.0 / (k + rank + 1)
chunks[chunk.id] = chunk
ordered = sorted(scores.items(), key=lambda kv: kv[1], reverse=True)
return [chunks[cid] for cid, _ in ordered[:limit]]
def retrieve(query: str, principal: Principal, top_k: int = 8):
# Permission filter is applied INSIDE each retriever, pre-ranking.
acl = principal.group_ids
dense = vector_index.search(embed(query), limit=50, filter={"acl_any": acl})
lexical = bm25_index.search(query, limit=50, filter={"acl_any": acl})
candidates = reciprocal_rank_fusion([dense, lexical], limit=50)
# Cross-encoder reranking: slow per pair, but only 50 pairs.
scored = cross_encoder.rank(query, candidates)
keep = [c for c in scored if c.score >= RELEVANCE_FLOOR][:top_k]
if not keep:
raise NoGroundingAvailable(query) # refuse, do not improvise
return keep
Retrieve wide, then rerank
The most reliable quality improvement available in a RAG pipeline is a cross-encoder reranker, and it is underused because it looks expensive.
The reason it works is architectural. An embedding model encodes the query and the document independently — it never sees them together, so it cannot reason about how they relate. A cross-encoder reads the query and the passage jointly and scores actual relevance. It is far too slow to run over your whole index, which is why you use approximate nearest neighbour search to get fifty cheap candidates and then spend real computation ranking only those fifty.
Trusting the top five straight from vector search is leaving most of your quality on the table. Retrieving fifty and reranking to eight typically adds tens of milliseconds — an order of magnitude less than the generation call you are about to make — and it is the difference between the right passage being at position two and being at position nineteen where it never reaches the model.
More context is not better
Large context windows tempt teams into sending fifty chunks and letting the model sort it out. This degrades answers, and it does so in a way that is easy to miss in casual testing.
Attention over long contexts is not uniform. Material at the beginning and end of the window is used far more reliably than material in the middle — the lost-in-the-middle effect is well documented and it is not a bug you can prompt your way out of. Padding the window with twenty marginally relevant chunks lowers the probability that the model uses the one chunk that mattered, because you have buried it.
Past a fairly low threshold, precision beats recall. Set a relevance floor and send fewer, better passages. Order them with the highest-scoring first. And measure this: take a question set where you know the answer passage, and compare answer quality at five chunks versus twenty. The result is usually the opposite of the intuition.
Permissions are a retrieval concern, not a prompt concern
This is the most serious production mistake we see, and it is worth stating bluntly: instructing the model not to reveal documents the user is not allowed to see is not access control. It is a request. The content is already in the context window, the prompt is not a trust boundary, and the failure mode is a compliance incident rather than a bad answer.
Filter at retrieval time, inside the query to the index, as a pre-filter rather than a post-filter. Post-filtering after ranking is also wrong for a subtler reason: if you retrieve the top fifty globally and then drop the ones the user cannot see, you may be left with three, and you have silently degraded the answer for privileged content the user could legitimately have received further down the list.
Note also that permission checks must be evaluated at query time against current group membership. Baking permissions into the index at ingestion means every access-control change requires a reindex, and until that reindex runs, your index is wrong. Building this properly is ordinary, careful platform and data engineering — it is the part of a RAG system that looks least like AI work and carries the most risk.
Grounding, citation and the ability to refuse
A production system must be able to say it does not know. This sounds obvious and is routinely omitted, because a system that answers everything demos better than one that declines.
Three mechanisms make refusal real. First, a relevance floor: if the best reranked passage scores below threshold, do not call the generator at all — you already know the answer is not in the corpus. Second, an explicit instruction that answers must be supported by the supplied passages, with a defined output for the unsupported case. Third, and most important, citation enforcement: require the model to attach passage identifiers to each claim, then verify programmatically that every cited identifier was actually in the context you sent. Claims without a valid citation get stripped or the response is regenerated.
Citations serve two purposes and the second is the bigger one. They let a user verify an answer, and they make the system's errors visible instead of invisible. A wrong answer with a citation is debuggable. A wrong answer without one is indistinguishable from a right one.
You cannot improve what you do not measure
Most teams evaluate by asking a few questions and forming an impression. That does not survive the first tuning change, because you have no way to know whether a change that fixed three questions broke nine others.
Build a golden set — 100 to 200 real questions, drawn from what users actually ask, each labelled with the passage or passages that contain the answer. This is a few days of unglamorous work and it is the highest-return investment in the entire project.
Then measure retrieval and generation separately, because they fail separately and fixing one does nothing for the other.
| Symptom | Likely stage | What to change |
|---|---|---|
| Answer invents plausible detail | Retrieval — answer passage absent | Hybrid search, reranking, chunk boundaries |
| Fails on error codes, part numbers, names | Retrieval — dense-only | Add BM25 and fuse |
| Right document, wrong or partial passage | Chunking | Structure-aware splits, parent expansion |
| Correct passage retrieved, answer still wrong | Generation or context order | Fewer chunks, better ordering, tighter prompt |
| Confidently quotes superseded policy | Metadata | Version and effective-date filters |
| Quality collapsed after a deploy | Operations | Embedding model version changed under you |
For retrieval, recall@k is the number that matters: in what fraction of questions does the correct passage appear in the top k you send to the model? It is a hard ceiling — if recall@8 is 0.62, then 38% of your questions cannot be answered correctly no matter how good the generator is. Track MRR alongside it to see whether the right passage is arriving near the top or scraping in at the bottom.
For generation, measure faithfulness (is every claim supported by the retrieved passages) and answer relevance (does it address what was asked). An LLM judge is a reasonable instrument here, with two caveats: judges are biased toward longer and more confident answers, and they must be validated against human labels on a sample before you trust their verdicts. Use them for regression detection, not for absolute quality claims.
Operating the thing
Three operational realities catch teams out, and all three are cheap to handle if you plan for them and expensive if you do not.
Pin your embedding model version. Vectors from different model versions are not comparable. Silently upgrading the embedding model means your query vectors live in a different space from your indexed vectors, and quality degrades in a way that looks like a mysterious regression. Treat the embedding model as part of the index identity, build a new index when it changes, and cut over deliberately.
Plan incremental updates from day one. A full rebuild is fine at ten thousand documents and unacceptable at ten million. You need change detection, per-document upsert and delete, and a way to reconcile the index against the source of truth — deletions are the ones that get forgotten, and an index that still serves a retracted document is a real problem.
Know your latency budget before you design. Embedding, hybrid retrieval, reranking and generation each take a slice. Reranking is usually worth its cost; generation dominates. If the budget is tight, cut chunk count before you cut the reranker. Cost and latency in the generation stage have their own set of levers, which we cover in our note on controlling LLM cost and latency in production.
What this means in practice
Build the evaluation harness before you tune anything. Without a golden set you are not engineering, you are reacting to anecdotes, and every change you make will be an uncontrolled experiment.
Then work in the order the failures occur: chunking and metadata first, because nothing downstream can recover a destroyed passage; hybrid retrieval and reranking second, because that is where the largest measurable gain sits; grounding and refusal third; prompt tuning last, because it is where teams instinctively start and where the least value is. Put permission filtering into the retrieval layer at the beginning — retrofitting a security boundary into a shipped system is the expensive version of this lesson.
None of this is exotic. It is careful information-retrieval engineering with a language model at the end, and the teams that treat it that way get systems that hold up. If you are staring at a prototype that will not survive its corpus, describe the corpus to us — that is usually where the diagnosis starts, and it is the shape of most of our AI engineering work.