EigenForge AI Labs Talk to us

Blog

A citation is not evidence

A link that resolves proves a source exists, not that it says what the sentence claims. Here is why the gap matters and how to measure it.

← All articles

2026-10-04 · EigenForge AI Labs

A footnote does two jobs in human writing. It tells you a source exists, and it makes a promise that the source supports the sentence it is attached to. Readers rarely check the second job. They assume that anyone who bothered to attach a source has read it.

That assumption is where AI-generated research goes wrong. A language model can attach a reference to every sentence in a report at almost no cost, and the references look exactly like the ones a diligent analyst would have added. The first job is done every time. The second is done some of the time, and nothing on the page tells you which times.

This essay argues three things: that much cited AI output would fail a simple check, that the check is not hard to describe, and that anyone relying on such output should run it before trusting the formatting.

Three things a citation can mean

When a report says "(Source 4)", it can mean three different things, and they are easy to confuse.

The first is that the source exists. The link resolves, the paper is real, the page loads. This is the property people check when they check anything, because it is cheap, and it is the property models are best at.

The second is that the source is about the topic. A retrieval system finds documents that are semantically close to the question, so the cited page is usually on the right subject. A page on supplier concentration risk will be cited for a sentence about supplier concentration risk.

The third is that the source supports the claim: the specific proposition in the sentence, with its numbers, its scope, its date and its level of certainty, follows from what the source says. This is the only property that makes a citation evidence, and it is the one that is rarely tested.

A source can pass the first test and fail the second, and pass the second and fail the third. The distance between the second and the third is where most of the damage sits, because a topically relevant source gives a sentence the air of having been checked.

A CITATION IS NOT EVIDENCEWhole-answer citationCommon, cheap, weakClaim-level attributionWhat an audit actually needsClaim 1Claim 2Claim 3Five sourceslisted at the endClaim 2 may be unsupported. Nothing reveals that.Claim 1Passage 1, record idClaim 2Passage 2, record idClaim 3Passage 3, record idA claim with no supporting span is refused, not softened.Doing this inside a budget an enterprise will pay for is not solved. It is one of our open questions.
Exhibit 02Whole-answer citation against claim-level attribution.

What current research is finding

Work on this has accumulated quickly over the past year, and it points the same way.

A 2026 preprint, Cited but Not Verified, evaluated citations from fourteen language models acting as deep research agents. It scored each inline citation on three dimensions: whether the link works, whether the content is relevant, and whether the fact stated is accurate against the source. The authors report that even the strongest frontier models keep link validity above 94 per cent and relevance above 80 per cent, yet reach only 39 to 77 per cent factual accuracy. They also report that factual accuracy fell by roughly 42 per cent on average, across two frontier models, as the number of tool calls grew from 2 to 150. More searching did not produce better attribution. It produced more material to attribute carelessly.

An earlier audit, DeepTRACE, broke the problem into measurable dimensions at the level of the individual statement. Across the generative search engines and research agents it examined, the fraction of citations that accurately reflected what the source said ranged from about 40 to about 80 per cent, and the share of statements with no support in the system's own listed sources varied enormously from one system to the next. Its authors also conclude that more sources and longer answers do not guarantee better grounding.

A third paper, From Fluent to Verifiable, argues for what it calls claim-level auditability. Its sharpest point is a definition. Soundness means that cited sources entail the attributed claims, and not merely that citations resolve to real papers.

These are preprints. Their numbers come from particular benchmarks, particular query sets and, in part, from models acting as judges, and a different query set would move them. Nobody should read "39 to 77" as a statement about their own system. The pattern is what carries over: surface metrics look healthy while the property that matters is much weaker, and it varies widely between systems.

Why a real source can fail to support the sentence

It helps to know the mechanisms, because they tell you what to test for.

The model often writes the sentence first and attaches a source second. Generation and attribution are separate steps, and the second step searches for something plausible instead of something entailing. Retrieval measures topical similarity, and topical similarity is not entailment.

Sources get compressed. A study says an intervention was "associated with" a lower rate in a particular group over a particular period, and the sentence says the intervention reduces the rate. Every word of the second sentence sits in the neighbourhood of the first, and the claim is stronger.

Scope slips. The source covers one jurisdiction and the sentence is general. The source is from 2021 and the sentence is in the present tense. The source reports a median and the sentence says average.

Claims get borrowed. The page that was cited quotes another report for the figure, and the model attributes the figure to the page. The number may be right and the provenance still wrong.

And some sources contradict the sentence. A page explaining why a common belief is mistaken will be retrieved for the belief, because that is its topic.

None of this needs the model to be careless in any human sense. It follows from how the pipeline is built, and it gets no better when the pipeline is given more material to work with.

THE COMMON APPROACHChunk, embed, retrieve bysimilarity✕Parent–child relationships are lost✕Joins cannot be rebuilt from similarity✕The copy drifts from the source✕Fluent, confident, sometimes wrong✕Content leaves the boundary to be embeddedAN ONTOLOGY OVER LIVE DATAModel relationships, resolve,query✓Table relationships preserved, joins correct✓Answers generated from live records, never from a stale copy✓Every answer traces back to records✓Policy applied at query time, per user✓Runs against a local model — nothing leavesIs it answering from a copy of your data, or from your data?
Exhibit 01Where the two approaches diverge, and the question to put to a vendor.

What a real check looks like

The check is a procedure, and it can be written down in about a paragraph.

Break the output into atomic claims, each one a single proposition. A sentence with three figures contains at least three claims. This is tedious by hand and easy to automate, which is part of why it is usually skipped.

For each claim, retrieve the passage the citation points to. The unit has to be a span of text, a few sentences, not a whole document. A citation that points to a hundred-page report cannot be checked in reasonable time, so it cannot be relied on, whatever the report says.

Then judge the claim against the passage, with four outcomes instead of two: supported, partly supported, not supported, contradicted. The middle outcomes matter. Most failures are "partly", and a pass or fail check loses them.

Check scope separately and mechanically. Do the dates, the population, the quantifiers and the numbers match exactly? A transposed figure, or an "all" that should be "most", is among the commonest failures and the easiest to catch if someone is looking for it specifically.

Look for conflict. Did retrieval turn up other material that disagrees? An answer that reports one side of a disputed question, with citations that all lean the same way, passes a claim-by-claim check and still misleads. The auditability paper calls this contradiction transparency, and it is a property that systems are unlikely to have unless somebody built it in.

Use a different judge from the generator. A model checking its own work shares its own blind spots. A second model is better, and a person sampling the results is better again.

Finally, keep the record: which claim, which passage, which verdict, which reviewer, which version of the answer. If the check leaves no trail, it cannot be repeated, and you cannot tell whether things are improving.

You do not have to check everything. Sample. Draw a set of answers at random, check every claim in them, and you have a rate you can track over time. The auditability paper makes a useful demand here, that audit effort should stay well below generation effort. A check that costs more than writing the report will not survive its first deadline.

The counter-arguments

There are three fair objections.

First, a model judging entailment is itself fallible. That is true. It is why the judge should differ from the generator, why people should sample, and why the rate you measure is an estimate with error bars, not a fact. It is still far better than no measurement, which is the default in most organisations.

Second, not every sentence maps to a single passage. A good analysis synthesises, and its conclusion follows from several sources together. Also true. The answer is to make the distinction visible. A sentence that restates a source should carry a span. A sentence that is inference should say so and name the claims it rests on. The failure is not synthesis. It is synthesis dressed as quotation.

Third, a strict check makes a system decline too often, and a tool that declines every second question will be abandoned. This is a genuine tension and it is a design decision to be set per use case. Where a wrong answer costs an awkward conversation, a looser threshold is reasonable. Where it costs a regulatory filing, it is not. What cannot be defended is having no threshold at all, and discovering your real one when someone outside the organisation finds the error.

What to ask of any system that cites

If you are buying or building one, four questions separate attribution from decoration.

Does each citation point to a passage, or only to a document? Can the system decline to answer when the retrieved material does not support a claim, and does it say why? Is verification done by a different process from generation? And when you ask it a month later to show its working, does the record exist?

If the answers are no, the citations are formatting. Formatting has some use, since it tells a reader where to start looking, but it should be described for what it is.

A test you can run this afternoon

Take ten answers your current system has produced, preferably ones that somebody relied on. For every cited claim, open the source, find the sentence that supports it, and highlight it. Where you cannot find one, mark the claim. Then count.

The count will not be a benchmark. It will be something more useful: a figure about your own material, produced by your own people, that you can bring back in a month and compare. We build grounded-retrieval systems for a living, and our own bias is towards systems that decline when the evidence is thin. Apply the same test to ours.

groundingevaluation

Disagree with this?

These are written to be argued with. If you think this is wrong, we would rather hear it than not.

hello@eigenforgelabs.ai

Send opens your email client with the note already addressed to us — nothing is stored on this site, and the message goes from your own mailbox, so our reply lands in yours.