Retrieval demos are dishonest in a very particular way. Not deliberately — it's just that the corpus in the demo is a hundred clean documents, all current, all in one format, all about different things.
A real corpus is ten years of files. Eight near-identical versions of the same policy. Scanned PDFs where the text layer is a rumour. A naming convention that changed twice. Documents that contradict each other, both of which are technically still in force.
Embed that and search it and you get answers that are fluent, confident, and wrong. Here are the four failure modes, in the order they show up.
Failure 1: Semantic similarity isn't what you asked for
Vector search finds text that means something similar. Often you need text that contains a specific token — a reference number, a company UEN, a clause identifier, a surname.
Ask for "clause 7.3" and embeddings will happily return clause 7.2, clause 8.1, and a paragraph about clauses in general. All semantically adjacent. All useless. The one thing an exact-match keyword index would have nailed in a millisecond.
Fix: hybrid search. Run lexical (BM25-style) and vector retrieval in parallel and fuse the results. This isn't a refinement, it's table stakes. Any domain with identifiers — legal, compliance, finance, logistics, healthcare — will fail on pure vector search, and it fails silently, which is worse.
Failure 2: Chunking destroys the thing that made the document readable
The default approach splits text every N characters with a little overlap. It's fast and it's fine on prose.
On real documents it's a shredder. A table's header row lands in one chunk and its numbers in another, so the retrieved fragment is a grid of figures with no idea what they measure. A clause gets severed from the definition it depends on. A row in a schedule loses the "unless subsection (b) applies" that inverts its meaning.
Fix: chunk along the document's own structure. Sections, headings, table boundaries, list groups. Keep a table intact even when it's oversized, and carry the heading path into every chunk so a fragment knows it came from Fees → Late payment → Corporate clients. That breadcrumb alone eliminates a whole class of confident wrong answers, because the model can see when the fragment is from the wrong section.
For scanned material, budget properly for extraction. A bad OCR pass poisons everything downstream, and no amount of clever retrieval recovers from text that was never read correctly. This is the least interesting part of the build and often the highest-leverage.
Failure 3: Retrieval is a shortlist, not an answer
Top-k vector search optimises for recall. You get ten plausible passages, of which perhaps two are actually responsive. Hand all ten to the model and you've made the answer worse: the eight irrelevant ones are context the model must actively ignore, and it doesn't always.
Fix: rerank. Take a generous candidate set — twenty, fifty — and run a cross-encoder that scores each candidate against the actual query rather than comparing pre-computed vectors. It's slower per document and dramatically more accurate, which is exactly the right trade when you're only scoring a shortlist.
Then cut hard. Five well-chosen passages beat twenty hedged ones, on quality and on cost.
Failure 4: Nothing knows what's current
The one that ends careers. Your corpus contains the 2021 policy, the 2023 amendment, and last month's revision. All three embed beautifully. Semantic search has no concept of "superseded" — it just knows all three are about the same topic.
So the agent cites the 2021 version, in a professional context, to a client. The answer is fluent and well-sourced and completely out of date.
Fix: treat freshness as data, not vibes. Every chunk carries an effective date, a version, and a supersession pointer. Retrieval filters on currency by default. And when superseded material is returned — sometimes you genuinely need the version in force at the time of an event — the model is told explicitly that it's historical.
Then make it visible. Every answer shows its sources with their dates. A user glancing at "source: Policy v2, effective March 2021" catches in one second what an eval suite might miss for a month.
The part that actually decides the project
None of the above is exotic. Hybrid search, structure-aware chunking, reranking, and freshness metadata are known techniques, and they're most of the difference between a retrieval system that survives contact with a real archive and one that doesn't.
What's harder is the discipline underneath: you cannot improve retrieval you don't measure. Retrieval quality is separately measurable from answer quality, and it should be. Build a set of real queries with the passages that should come back, and score retrieval on its own. When the agent gives a bad answer, you then know immediately whether the model reasoned poorly or was simply handed the wrong documents.
Nine times out of ten it was handed the wrong documents. Which is good news — that's the half you can fix without touching a model at all.
Friday: three prompts we deleted, and what replaced them.
