EigenForge AI Labs Talk to us

Blog

When a knowledge graph is the wrong answer

Graph retrieval is over-applied. Here are the cases where plain vector search wins, where a graph is over-engineering, and the criteria that should decide.

← All articles

2026-11-01 · EigenForge AI Labs

A knowledge graph is an expensive way to answer questions that a search index answers adequately. That sentence will annoy people who have spent a year building a graph, and it should be said plainly, because a great deal of current enterprise AI work starts from the assumption that a graph is the sophisticated option and a vector index is the naive one. For many workloads that has the order backwards.

The evidence is more mixed than either camp admits. The benchmark paper When to use Graphs in RAG opens with the observation that graph-based retrieval "frequently underperforms vanilla RAG on many real-world tasks". The systematic comparison RAG vs. GraphRAG finds that plain retrieval does well on single-hop and detail-oriented questions, while graph methods do better on multi-hop questions and on broad summaries. The second also finds that routing each question to the better method, or combining both sets of results, beats either alone. A practitioner summary in VentureBeat, Stop graphing everything, reaches the same recommendation and adds a caution: many of the published gains rely on language models acting as judges, and those judges have biases that can swing results.

So the question is not which technology is better. It is which questions you have, and what each approach costs you to answer them.

What vector retrieval does well

Vector retrieval, usually paired with keyword search and a reranking step, finds passages that resemble the question. If the answer lives in one passage, or can be assembled from two or three adjacent ones, this is close to ideal. A policy lookup, a clause in a contract, the specification of a part, the steps of a procedure: all of these are questions where the answer is written down somewhere, in words similar to the question's, and the job is to find that place.

It also has virtues that are easy to undervalue. It is cheap to build and cheap to refresh. When a document changes you re-index that document. It makes no early commitment about what the world contains, so a new kind of document, or a question nobody anticipated, still works. And it retrieves original text, which is what you want to cite. A graph built by an extraction model is a paraphrase of the source, and a paraphrase is one step further from the evidence.

On the questions that dominate most internal assistants, which are lookups of one kind or another, the benchmarks above suggest that a well-tuned plain pipeline is hard to beat, and sometimes cannot be beaten. The first thing to build is a strong baseline: hybrid search, sensible chunking that respects tables and headings, a reranker, and a model instructed to decline. Most teams that skip this step and go straight to a graph have no way of knowing what the graph bought them.

Where a graph is over-engineering

Several patterns recur.

The graph was built because the diagram looks impressive. A node-and-edge picture is an effective thing to show a steering committee, and it can substitute for a clear statement of which questions the system must answer. If nobody can name the three hardest questions the graph exists to handle, it is probably decoration.

The graph is a vector store with entity tags. A good deal of what is sold as graph retrieval amounts to retrieving chunks and then decorating them with a few extracted entities. That may help, but it does not need the machinery, and it does not deliver what the word graph suggests.

The graph is extracted from prose by a language model and never verified. Extraction is lossy and sometimes wrong, and a wrong edge is silent. A wrong chunk is usually visible when read, because it is just text. A wrong relationship between two entities sits inside a structure that looks authoritative, and nobody sees it until an answer depends on it. If you build a graph this way, you need a plan for sampling edges and measuring their accuracy, and many projects have none.

The ontology has no owner. Entity types and relationships are decisions about how the organisation sees itself, and they change when the organisation does. A schema that is not maintained turns into a map of how the business worked at the time of the project. Maintenance is the real cost of a graph, and it is the one least often budgeted.

The corpus is small, or fast-changing, or both. Where there are a few hundred documents, almost any retrieval method works, and a graph adds cost for no visible gain. Where documents change daily, re-extraction becomes a standing expense, and the graph is always slightly out of date.

The indexing bill is real. Some graph methods use a language model to process the whole corpus before a single question is asked, and the cost is paid again whenever material changes. The benchmark paper above reports substantial differences between graph methods in the tokens they consume, and notes that excess retrieved material can reduce the relevance of the context given to the model. Cheaper graph variants exist and are improving, so this argument weakens with time, but it has not gone away.

THE COMMON APPROACHChunk, embed, retrieve bysimilarity✕Parent–child relationships are lost✕Joins cannot be rebuilt from similarity✕The copy drifts from the source✕Fluent, confident, sometimes wrong✕Content leaves the boundary to be embeddedAN ONTOLOGY OVER LIVE DATAModel relationships, resolve,query✓Table relationships preserved, joins correct✓Answers generated from live records, never from a stale copy✓Every answer traces back to records✓Policy applied at query time, per user✓Runs against a local model — nothing leavesIs it answering from a copy of your data, or from your data?
Exhibit 02The honest comparison, including where the graph loses.

Where a graph earns its keep

There are questions that plain retrieval cannot answer well, and it helps to be exact about them.

Questions that require joining records. Which supplier's batch went into which product, and which customers received it? The answer is not in any single passage. It exists as a path through several records, and a similarity search will not follow it. This is the multi-hop case, and it is where both papers find graph methods ahead.

Questions that require counting or enumerating. How many contracts expire this quarter, and which ones? Retrieval returns the top few passages, not all of them, and a language model asked to count from a handful of fragments can produce a confident number that is wrong. Enumeration needs a structure you can query exhaustively.

Questions about the whole corpus. What are the main themes across these ten thousand reports? A top-k retrieval sees a sliver. Graph-based methods that summarise communities of related entities were designed for exactly this kind of global question, and they do better on it, with the caveat, noted above, that many of the published comparisons use a language model as judge.

Questions where identity is the problem. The same customer appears under three names in three systems, and a question about "this customer" needs them resolved into one. Embeddings place similar names near each other, but they do not decide that two records are the same entity.

Settings where relationships carry the permission or lineage. If who may see what depends on how entities relate, or an answer has to be traced through a chain of derivations, structure is not optional.

Questions where the explanation matters as much as the answer. A path through a graph, from one record to the next, is a natural thing to show a reviewer, and it is easier to audit than a block of retrieved text that happens to contain the right words. This is a genuine advantage, and it holds only if the edges are verified. A traceable path through unverified edges is a well-presented guess.

A last consideration cuts both ways. A graph commits you early. Choosing entity types and relationships means deciding what the organisation considers worth representing, and anything outside that scheme is invisible to the graph. For stable, well-understood domains that discipline is a benefit, because it forces agreement on definitions that people were previously arguing about. For a domain still being understood, it can freeze a premature view of the world into the index.

There is a distinction here that the debate usually flattens. A graph extracted by a model from prose and an ontology defined over structured systems of record are different things. The first is a retrieval trick, and its value is whatever the benchmarks say it is. The second is a data model: records you already hold, with relationships that were defined and validated by the people responsible for them. Many of the questions in the list above are of the second kind, and many of them are, at bottom, well-posed queries over structured data. Whether you call the result a graph, a semantic layer or a set of governed joins is secondary. Sometimes the correct answer to "should we build a graph?" is "you need SQL and clean keys".

Criteria that should decide

Six questions do most of the work.

Does the answer live in one passage or across several records? If one passage, use plain retrieval. Is the question a lookup, or does it ask for aggregation, enumeration or a path? Lookups belong to vectors, the others to structure. Is your source material prose or structured data? Prose favours chunks, and structured data favours a model of the entities, built from the data and not extracted from text about it. Who owns the schema, and what happens when it changes? If the answer is "nobody", do not start. What does a wrong edge cost, and how would you notice one? If you cannot answer the second part, you are not ready to rely on the graph. And have you beaten a good baseline on your own questions? If you have not tried, you do not know.

The practical procedure follows from the last question. Collect the thirty or so hardest real questions your users ask, plus the ones they have given up asking. Run them through a tuned hybrid baseline. If the baseline answers most of them acceptably, stop, and spend the money on corpus quality. If it fails in a pattern, and the pattern is joins, counts or whole-corpus questions, you have found the specific place a graph should go. Build it for that, and route those questions to it. Leave the rest on the baseline. The routing and integration results in the comparison paper suggest this hybrid will beat either approach alone.

01Foundation & governanceone set of numbers · lineage ·cataloguing · migration assurance02Operations & productionyield · asset performance ·condition-based maintenance03Quality & compliancegenealogy · long-horizon retrieval ·containment scoping04Commercial & revenuecost to serve · counterparty 360 ·revenue assurance05Finance & controlconsolidation · close acceleration ·working capital06Intelligence & AIself-service · insight agents ·document intelligence07Governed agentic operationsthreshold-bound action · approvalrouting · sealed record08Sensing & field operationsdetection with a coordinate · offlinecapture · twinsEight groups. Fifty-plus patterns. Read it as range rather than a menu:if a question can be answered from data you already hold, it can be built.
Exhibit 01The range of questions each approach serves well.

A disclosure

We should say where we stand. EigenForge sells, among other things, retrieval grounded in an ontology built over live structured data, so we have an interest in graphs being taken seriously, and you should weigh this essay with that in mind. The discipline we try to apply is the one described above: start with the questions, build the baseline, and add structure only where the baseline fails in a way that structure fixes. If a client's questions are lookups over a document set, we would tell them to build a good search index and leave it there.

That advice does cost us, and it is still right. A graph earns its place by answering questions that a cheaper method cannot, and a team that can name those questions will build a better one than a team that started from the technology. The measure of success is whether the questions got answered, and the shape of the data structure behind them is only a means to that.

retrievalknowledge-graphsarchitecture

Disagree with this?

These are written to be argued with. If you think this is wrong, we would rather hear it than not.

hello@eigenforgelabs.ai

Send opens your email client with the note already addressed to us — nothing is stored on this site, and the message goes from your own mailbox, so our reply lands in yours.