A RAG prototype takes an afternoon. Load documents, embed them, stuff the top five chunks into a prompt, answer questions. It demos beautifully, which is exactly the problem — the demo gives you no information about whether the production version will work, because every hard part of retrieval shows up only at scale, with real documents, in front of users who have consequences attached to the answers.
Here is what we have learned shipping these into regulated and operational environments, roughly in the order the problems appear.
Your ingestion pipeline is the system
Teams spend weeks on embedding models and reranker choices and an afternoon on parsing. The ratio should be inverted. We have never seen a production retrieval system where the dominant error source was the embedding model. It is almost always the pipeline that turned documents into text.
Specific things that silently destroy retrieval quality:
Tables flattened into prose. A financial statement parsed naively becomes a stream of numbers with no column association. The retriever may find the right chunk and the model will still get the answer wrong, because the chunk no longer contains the relationship between the number and its label. Use a layout-aware parser and keep table structure — markdown or HTML tables both work.
Multi-column PDFs read in the wrong order. Reading order on a two-column page is a layout problem, not a text problem. Naive extraction interleaves the columns and produces sentences that no longer mean anything. This is common in academic, legal and insurance documents.
OCR failing quietly. The worst failure mode in the pipeline, because it reports success. A scanned page at a slight angle yields plausible-looking garbage, which gets embedded, indexed, retrieved and cited — and the citation points at a real page, so it even looks correct in review. Gate on OCR confidence and route low-confidence pages to a human queue. On one engagement, 3% of pages failed this gate and those 3% accounted for a disproportionate share of early errors.
Headers and footers in every chunk. Repeated boilerplate adds embedding noise to every vector and can make near-duplicate chunks indistinguishable. Strip it during ingestion.
A practical test: sample fifty pages across the weirdest documents you have, extract them, and read the output yourself. Not a metric — read it. You will find things no metric would have surfaced.
Fixed-size chunking is a convenience, not a design
Splitting at 512 tokens with 50 overlap is the default everywhere and it is wrong for most real corpora because documents have structure that the token counter cannot see.
What works better: chunk along the document’s own boundaries. A contract clause is a chunk. A table is a chunk, whole. A policy section with its subsections stays together if it fits. Then — and this matters more than the chunk size — prepend context to each chunk: document title, section path, effective date, and a one-line summary of what the chunk is. A chunk that reads “the limit shall be $2,000,000” is nearly useless in isolation; “Acme Master Services Agreement 2024 → §8.3 Limitation of Liability → the limit shall be $2,000,000” retrieves and answers correctly.
Also: index at more than one granularity. Retrieve small chunks for precision, then expand to the parent section before generation. Small-to-large retrieval is cheap to implement and routinely the single biggest accuracy gain available after fixing parsing.
Dense-only retrieval fails on exactly the queries that matter
Vector search is excellent at paraphrase and conceptual similarity. It is unreliable at exact match — and in business documents, the high-stakes queries are nearly all exact match: an account number, a specific entity name, a part number, a statute reference, a defined term.
Embeddings place “Acme Holdings LLC” and “Acme Holding Company Ltd.” close together. In a credit file, those are different legal entities and conflating them is a material error. BM25 does not make that mistake.
Run both and fuse. Reciprocal-rank fusion needs about fifteen lines of code, requires no tuning, and in our benchmarks recovers most of the gap that dense-only retrieval leaves on entity lookups. Then rerank the fused candidates with a cross-encoder — retrieve 50, rerank to 8. The reranker is usually the best accuracy-per-dollar component in the stack, because it sees the query and the passage together rather than comparing two independently-computed vectors.
A system that cannot say “I don’t know” is not deployable
This is the difference between a demo and a product, and it is a design decision rather than a prompt.
Instruct a model to answer from context and it will answer from context — including when the context does not contain the answer. It will assemble something plausible from adjacent material, and plausible wrong answers are far more dangerous than refusals because they survive review.
What actually works:
- Threshold on reranker scores, not similarity scores. Cosine similarity is not calibrated and the absolute value means little. Cross-encoder scores are usable as a confidence signal.
- Make abstention a first-class output. A schema field, validated, not a hoped-for phrase in free text.
- Measure abstention separately. Track answer accuracy on answered questions, abstention rate, and false-abstention rate as three distinct numbers. A single accuracy figure hides the tradeoff you most need to see.
- Route abstentions somewhere useful. Hand the human the retrieved candidates. An abstention with context attached still saves most of the reading time.
On a document-review platform we built, the system abstains on about 11% of questions. The client regards that as a feature — it is the number that let internal audit sign off.
Citations must be enforced, not requested
“Cite your sources” in a system prompt produces citations most of the time. Most of the time is not a control.
Enforce it structurally: require citations in the output schema, validate that each cited id resolves to a chunk that was actually retrieved for this query, and reject-and-retry when it does not. Then escalate to a human after a retry. The model cannot emit an uncited claim to a user because the validator sits in between.
Resolve citations to something a human can verify in one click — document, page, highlighted region. A chunk id is not a citation to an auditor. In regulated work this detail is often the deciding factor in whether the system is usable at all, independent of accuracy.
Build the evaluation set before you build the system
The uncomfortable discipline that separates projects that ship from projects that linger.
Before writing retrieval code, have domain experts answer 100–150 representative questions against real documents, with citations, independently of each other. This gives you three things:
- A gold set for measuring every subsequent change.
- A human consistency baseline. Experts typically agree 85–93% of the time. That number is your realistic ceiling, and it stops the conversation where someone asks for 99%.
- Discovery of ambiguous questions. Where experts disagree, the question is usually underspecified. Fixing the question is free accuracy — on one project, six of 31 checklist questions were ambiguous and rewording them improved measured performance before any model work.
Measure retrieval and generation separately. Recall@k and MRR for retrieval; faithfulness, answer correctness and citation accuracy for generation. When end-to-end accuracy drops you need to know which half moved. Teams that only measure end-to-end spend days bisecting by hand.
LLM-as-judge is workable for scale, but calibrate it against human adjudication on a sample each cycle and track the agreement rate. An uncalibrated judge drifts and you will not notice.
Production realities that surprise people
Cost is dominated by input tokens. Retrieved context, not generated answers. Retrieving 20 chunks instead of 5 roughly quadruples cost for typically marginal accuracy gain. Rerank aggressively and send fewer, better chunks.
Latency is dominated by reranking and generation. Vector search is milliseconds. Budget realistically and stream output — perceived latency matters more than actual for interactive use.
Index freshness is a product decision. How stale can an answer be? Minutes, or a day? That answer determines whether you need incremental indexing with change-data capture or a nightly rebuild. It is much cheaper to decide early.
Permissions must be enforced at retrieval time. Filter the candidate set by the requesting user’s access rights before the model sees anything. Post-hoc filtering of a generated answer leaks — the answer already contains the content. We have seen this designed wrong more than once.
Users will ask questions the corpus cannot answer. Analytical questions (“which of these contracts is riskiest?”) against a corpus indexed for lookup. Detect and redirect, or you accumulate a long tail of confident nonsense.
The short version
Spend your time on parsing, structural chunking, hybrid retrieval with reranking, enforced citations, calibrated abstention and an evaluation set built by domain experts before the first line of retrieval code. The embedding model matters least of anything on that list.
We have shipped retrieval systems into regulated, air-gapped and high-volume environments. Book a consultation, or read how this played out on a 4.1M-page credit-file platform.