EigenForge AI Labs Talk to us

Blog

Retrieval is not one thing

"RAG" now covers five meaningfully different architectures with different strengths, costs and failure modes. Choosing between them is an engineering decision — and the most common choice is wrong for enterprise questions.

← All articles

2026-10-05 · EigenForge AI Labs

Three years ago "RAG" meant one thing: chunk your documents, embed the chunks, retrieve the similar ones, and let the model answer from them. In 2026 the word covers at least five meaningfully different architectures, and procurement conversations that treat them as one product keep ending in the same place — a system that demos beautifully on the handbook and fails quietly on the questions the business actually asks.

The distinction that matters is not which embedding model. It is what kind of thing your questions are.

Five patterns, briefly

Naive RAG is the original shape: documents become vectors, similarity finds the candidates. It is fast to stand up and genuinely good at questions whose answer lives inside one document. Its failure mode is structural: the copy drifts from the source, and relationships between records — the thing enterprise answers are made of — do not survive chunking.

Advanced RAG keeps the shape and sharpens the parts: re-ranking, query rewriting, hypothetical-answer retrieval. It recovers a lot of recall on messy corpora. It remains a copy of the documents, ageing from the day it was built.

Modular RAG treats retrieve, re-rank and generate as separate, measurable stages. This is less a technique than a discipline — each stage can be evaluated and swapped independently — and it is the minimum seriousness for a system you intend to operate rather than demo.

Graph RAG retrieves structure instead of passages: entities, and the relationships between them, so a question that spans a customer, their contracts and their payment history is answered from the joins, not from three nearby paragraphs. Where it fits, nothing else comes close. Where it does not — and there are honest analyses arguing it is over-applied — it is an expensive way to answer questions a table would have handled.

Agentic RAG lets an agent plan retrieval as a sequence of calls: decompose the question, fetch, notice what is missing, fetch again. It handles the hardest, least-specified questions, at compound cost and latency — which is why it belongs behind a budget, on the question classes that justify it.

FIG. 1 RETRIEVAL IS NOT ONE THING — FIVE PATTERNS, GRADED HONESTLYPATTERNWHAT IT ISWHERE IT WINSWHERE IT FAILSUSE IT WHENNAIVE RAGchunks, embedded, retrieved bysimilarityfast to stand updrifts from source;relationships lostdocuments that stand aloneADVANCED RAGre-ranking, query rewriting, HyDEbetter recall on messycorporastill a copy, still ageingpolicy and handbook librariesMODULAR RAGswappable retrieve / rerank /generate stageseach stage measuredseparatelyintegration burden is yoursteams with evaluationdisciplineGRAPH RAGentities and relations retrieved asstructuremulti-hop questions,lineageover-applied where a tablewould doquestions that span recordsAGENTIC RAGan agent plans retrieval as asequence of callshard, underspecifiedquestionscost and latency compoundresearch-grade questions,budgetedThe last two rows are where enterprise questions actually live — anything whose answer spans more than one record. That is theretrieval problem our research programme works on.
Exhibit 01Five retrieval patterns, graded honestly — choose by the question, not the brochure.

The technical detail

Two engineering truths sit underneath the taxonomy.

The first: evaluation has to be per-pattern and per-question-class, on your questions. "Which RAG is best" is a leaderboard question. "Which retrieval architecture answers our reconciliation questions defensibly, at what cost per answer" is an engineering question, and it has a different answer for the policy library than for the ledger.

The second: the corpus problem and the engine problem look identical from outside. When answers are poor, either retrieval and generation are weak, or the answer simply is not in the material. These need opposite fixes, and teams routinely rebuild the engine when the corpus was the problem. Before changing any architecture, separate the two with a labelled question bank — fifty questions whose answers you know, annotated with where the answer lives. It is a week's work and it has saved our clients entire rebuilds.

Where we land

Most enterprise questions worth asking span more than one record. "Which suppliers are exposed to both sanctions risk and a contract renewal this quarter" is not a passage-retrieval question; it is a question about relationships over live records, with entitlement applied to the asker.

That is why our platform resolves questions against an ontology over records rather than a vector copy of documents — retrieval as query, not as search. We hold it as a position tested on every engagement, and we read what argues against it — the failure analyses of graph approaches included. But the direction of travel in the field — from passages toward structure, and from structure toward agents that plan — is the direction we started from.

Choose the pattern by the question, not the brochure.

ragretrievalarchitecture

Disagree with this?

These are written to be argued with. If you think this is wrong, we would rather hear it than not.

hello@eigenforgelabs.ai

Send opens your email client with the note already addressed to us — nothing is stored on this site, and the message goes from your own mailbox, so our reply lands in yours.