Why I stopped retrieving chunks at query time, and what I do instead

While testing our internal assistant I asked what a particular limit was. It said 30. I phrased it a bit differently and it said 100.

Both answers were confident and both cited real documents. The two numbers were in two different documents tho and one of them was two years old. The system had no idea, and no way to have an idea. Every time anyone asked, it went and picked again.

That was the day I stopped building on RAG for knowledge bases (true story)

RAG is fine. It’s just not what most people need.

Let me be fair before I start throwing punches.

If you have a big pile of documents and people ask unpredictable questions across all of it, retrieval-augmented generation is the right tool. Search over a corpus is a well-understood problem. Vector search does it well. Nothing I’m about to say makes that untrue.

But most people reaching for RAG are not building search.

They’re building a brain. A thing that holds what a team or a company or a research area actually knows, and answers questions from it in a way you can rely on. That’s a different job, and RAG is the wrong shape for it in ways that don’t show up in a demo and show up badly six months in.

Here’s what breaks.

1. It re-decides the same thing every single time

Two documents disagree. Both are in your index. Both are chunked, embedded, sitting there.

Someone asks. Retrieval pulls some subset of the relevant chunks, in whatever order the similarity scores landed. The model reads them and produces an answer. Maybe it picks one. Maybe it splits the difference. Maybe it goes with whichever appeared first in the context window.

The disagreement never gets resolved. It gets re-litigated on every query, non-deterministically, by a model that has no way to know which document is newer, which team owns it, or whether one of them was superseded by a decision made in a meeting last March.

This is the part that gets me. The system contains a genuine conflict, and instead of surfacing that once so a human can fix it, it hides it behind a fresh coin flip on every request.

2. Chunks don’t know what they are

A chunk is a slice of text with a vector next to it.

It doesn’t know it came from version 3 of a document that’s now on version 7. It doesn’t know the paragraph directly above it started with “in the legacy system.” It doesn’t know the doc was written by an intern in 2023 and never reviewed.

You can bolt metadata on. Everybody does. But the model isn’t reasoning over your metadata, it’s reading text that got amputated from its context and handed over with a confidence it hasn’t earned.

3. Nothing ever dies

Ingest a document. It’s in the index.

Six months later the document gets rewritten. You ingest it again. Now you have both versions in there, and retrieval has no concept that one of them is a corpse.

Ask most teams what their invalidation strategy is and the honest answer is “delete everything and re-index,” which nobody schedules and nobody does.

4. There’s no artifact

This is the one that actually made me stop.

Open your vector store. Go on. Show me what the system knows.

You can’t. There’s no object. There’s no file anyone can read, review, correct, hand to a new engineer, or diff against last month. The knowledge only exists as a temporary reconstruction inside a context window, and it evaporates the moment the request returns.

Every real knowledge system humans have ever built produced an artifact. A book, an encyclopedia, a wiki, a filing cabinet. Something you could point at. We built the first one that doesn’t, and then acted surprised when nobody trusts it.

5. When it’s wrong, you can’t tell why

Bad answer comes out. Now go debug it.

Was the knowledge missing entirely? Was it there but retrieval didn’t surface it? Did retrieval surface it and the model ignored it? Did it surface a stale chunk that outranked the current one?

You mostly cannot tell after the fact. And a failure mode you can’t diagnose is a failure mode you can’t fix, so it just stays.

The inversion: compile once, at ingest

Here’s the shape I use now.

Do the expensive thinking once, when knowledge enters the system, and make it produce something permanent.

On ingest: fetch the source. Convert it to clean markdown, strip the nav and the ads and the boilerplate. Hash the body, so the same content arriving twice costs you nothing at all. Then one model pass over the whole source — not chunks — that pulls out entities, checks for claims contradicting what’s already in the base, and decides what pages need to be written or updated. Render a real markdown page with structured frontmatter. Commit it to git.

On query: full-text search over the finished pages. Hand the model complete page bodies. Tell it to answer using only what it was given. Raw sources are never touched again.

The markdown files on disk are the source of truth. The database is a projection — an index, nothing more. If Postgres burns down you rebuild it by reading the files back in.

That one property changes how the whole thing feels to operate. Your backup strategy is git clone. Your disaster recovery is a folder.

What you get for it

Contradictions get settled when knowledge is written, not when it’s read.

Ingest finds a real conflict — same fact, different values — and the write stops. A human decides. Once. With context. And the decision lives in the page.

Softer conflicts get written but flagged, so they’re findable later without blocking anything. Cosmetic disagreements just get a marker in the file that a lint pass can scan for.

The point is that a query is never asked to arbitrate. Arbitration is a write-time concern.

You can actually read it.

The knowledge base is a folder of markdown. Anyone opens it. Anyone reviews it. Anyone corrects a page by hand and commits. Full history. When something’s wrong, you fix the page, and it’s fixed permanently for everyone — not just until the next retrieval rolls differently.

Refusal becomes a first-class outcome.

Make the model emit a specific token when it genuinely can’t answer from the pages it was handed. Then catch that token in code, not in vibes. When it fires: don’t cache it, don’t promote it, return an honest “I don’t know, here’s what we do have on this.”

A knowledge base that confidently makes things up is worse than no knowledge base, because the whole value proposition was that people would stop checking. If it lies even occasionally, they have to check everything, and you’ve spent six months building something with negative utility.

One rule that took me embarrassingly long to see: never cache an answer with no evidence attached. Your cache invalidation almost certainly works by matching updated pages against the evidence list of cached answers. An answer with an empty evidence list can never match anything, ever. It lives forever. Evidence-free answers must not be cached, full stop.

Cost moves to ingest.

You pay once per document instead of on every question. For a bounded corpus with repeat questions — which describes nearly every internal tool — that’s dramatically cheaper. Cache repeated questions on top and a big chunk of your traffic costs zero tokens.

What it costs you

I’m not selling this. It’s genuinely worse than RAG in specific, real ways, and you should know them before you build it.

Ingest is slow and expensive. A page can take a minute and real money. Bulk-loading ten thousand documents hurts.

It needs a human in the loop. Hard contradictions block until someone rules on them. If nobody watches the review queue, ingest quietly stalls and you find out three weeks later.

It’s bad at exhaustive retrieval. “Find every mention of X across four thousand files” is a search problem. Use search. This architecture answers questions from curated knowledge, it does not scan a corpus.

It only fits bounded, high-value collections. Internal documentation. A research area. One product’s knowledge. Not the open web.

So pick on this basis: is your failure mode missing something, or is it confidently saying something wrong?

Missing something, use RAG. Confidently wrong, come this way.

If you go build it

Things I learned the hard way, in no order.

Never truncate. When content exceeds context, chunk at paragraph boundaries and map-reduce the results back together. Truncation loses information silently, which is the worst possible way to lose information.

Postgres full-text search over a few hundred curated pages beats vector search, and more importantly you can debug why a query missed. Weight titles about ten times heavier than bodies. Run a strict AND pass first, and when that returns nothing, fall back to prefix matching ranked by how many query terms actually hit. If you’re bilingual, index in both languages and take the better rank.

Version your pages. When one updates, keep the old body as history rather than overwriting. Storage is free, context isn’t.

Serialise writes with an advisory lock. Single writer, concurrent reads. Kills a whole category of race conditions and you will never once miss the throughput.

Treat ingested raw sources as immutable. Retraction should be an explicit, loud, deliberate operation. Not a side effect of anything.

The gap I couldn’t close

Do all of that and you still can’t answer the question that matters most.

How much should I trust this page?

One page came in through a scraper that mangles tables, then a model that paraphrased it, then a template that filled a file. Another was written by hand by someone who knows the subject.

Both are markdown. Both have frontmatter. Both look exactly as authoritative as each other. Nothing in the architecture I just described distinguishes them.

That’s the problem I ended up publishing on.

Every page carries its transmission chain in frontmatter. Which component touched it, in what order, and whether that step was destructive, generative, or pass-through. Each component carries a reliability grade. The chain gets graded by its weakest link — because a great model sitting downstream of a broken scraper doesn’t rescue anything, it just launders the damage.

Then you cross the chain grade against a content-consistency check and you get an actual decision. Serve it. Serve it with a caveat. Route it to review. Quarantine it.

The method is adapted from isnād–rijāl in hadith science, which spent several centuries on precisely this problem: grade a report by grading the chain of people who transmitted it, and treat the weakest transmitter as the ceiling. Different substrate, same structure. Turns out the hard part was never the storage.

Paper: arxiv.org/abs/2607.24117 Code: pip install isnad — MIT, github.com/alizahidraja/isnad

It’s early. Someone I’ve never met filed four real bugs in it last month and I’d rather have more of that than less.

Take the pattern whether or not you take my library. If you’re building a brain on top of a knowledge base right now, the thing I’d most want you to walk away with is this: the expensive thinking belongs at write time, and whatever comes out of it should be a file a human can open.