Most AI pilots that stall do so quietly. Nobody cancels them. The demonstration impressed the sponsor, the first month of real use produced a pile of anecdotes about wrong answers, and the programme drifts into a state of being evaluated that lasts until the budget cycle ends.
The post-mortem, when there is one, usually reaches for one of two explanations. The first is that the model was not good enough: it made things up, it missed obvious points, it gave different answers on different days. The second, offered by whoever supplied the system, is that the data was not ready.
Both explanations can be true. Both are also routinely stated with no evidence at all, and that is the subject of this essay. A wrong answer from a retrieval-based assistant looks the same whether the material was missing, the material was found and misread, or the material was found, correct, and contradicted by another document the system also found. From the outside there is one symptom and at least three causes, and the fixes point in opposite directions. Replace the model when the corpus is the problem and you get the same wrong answer in better prose. Spend six months cleaning documents when the retrieval is the problem and the answers will not move.
Two systems, one symptom
Think of an assistant as two parts. The corpus is what the organisation has written down: policies, contracts, manuals, tickets, reports. The engine is everything that turns a question into an answer. That covers how documents are split and indexed, how passages are found, what the model is asked to do with them, and whether it is permitted to say that it cannot tell from this material.
A failure in either part produces a confident wrong answer, a vague answer, or an inconsistent one. Users report all three as "it's unreliable". The sponsor hears "the AI isn't ready". Nobody has located the fault, so nobody can fix it, and the next pilot repeats the first with a different logo on the slide.
A way to tell them apart
The test needs a few dozen real questions, one person who knows the answers, and an afternoon. It works by separating three steps that are normally blurred together: whether the answer exists, whether it was found, and whether it was used correctly.
For each question, start with the person, not the system. Ask them to find the answer in the corpus and note where it is. Sometimes they will not find it. Sometimes they will find two answers that disagree. Those outcomes are already corpus findings, and they were reached without running any AI.
Next, run the system and look at what it retrieved before reading what it said. Did the passage containing the answer appear among the passages the model was shown? If the answer exists and was not retrieved, the fault is in the engine's retrieval. Typical causes are chunking that separated a table from its heading, an index that matches words but misses the acronym your organisation uses, a permission filter that removed the document, or a scanned PDF that nobody extracted text from.
Then do the step that settles most arguments. Hand the model the correct passage yourself, the one your person found, and ask the question again. If it now answers correctly, generation is fine and the fault lies upstream of it. If it still answers wrongly, you have a generation fault: the model, the prompt, or the instruction about what to do when passages conflict.
Finally, ask questions the corpus cannot answer, and watch what happens. A system that answers anyway has an engine problem whatever the state of the documents, because it has no way to decline. It also makes every corpus problem look worse, since each gap in the documents becomes an invention in the output.
Lay the results out as a table of questions against outcomes and a pattern emerges that opinion cannot easily argue with. If most failures are "not in the corpus" or "contradictory", the corpus is the problem. If most are "in the corpus, not retrieved", retrieval is. If most are "retrieved, answered wrongly", it is generation. If most are "not in the corpus, answered anyway", the system lacks a refusal path. Most pilots turn out to be a mixture, and the proportions are the finding.
What corpus failure looks like
The commonest corpus failure is that the knowledge was never written down. It lives in the heads of three people who have been there for years, in a chat thread, in the margin of a printed manual. A pilot cannot answer questions that the documents cannot answer, and a surprising share of what an organisation believes it knows is of this kind.
The next is contradiction. There are three versions of the leave policy, none dated, and the authoritative one is a PDF of an email. Superseded procedures were never retired. Two departments use the same term for different things. A retrieval system will find all three versions and present one, or blend them, and the output is wrong in a way that no model upgrade can address.
Then there are documents nobody owns. Nobody can say whether they are current, so nobody can say whether an answer built from them is correct. The fix here is dull and mostly organisational: name an owner for each body of material, date it, say which version is authoritative, retire the rest, and write down what only people know. It cannot be bought, and it does not look like AI work, which is why it is often skipped in favour of something more exciting.
What engine failure looks like
Engine failures are more technical and, once located, more tractable. Retrieval that is measured only through the quality of final answers cannot be tuned, because an answer can be wrong for any of four reasons. The remedy is to measure retrieval on its own: for each test question, did the right passage come back, and in what position? Chunking, hybrid keyword and semantic search, handling of tables and scanned pages, and the treatment of your organisation's own vocabulary all show up in that number.
The model is the last thing to change, and it is usually the first thing changed. Swapping the model is the cheapest experiment to run, so it is run first, and because model output varies naturally from one run to the next, a swap often appears to help for a week. The diagnostic above says when it is the right move: only when handing the model the correct passage still produces a wrong answer.
Two causes outside the pair
Two other causes are common enough to name.
The question was never defined. The pilot was pitched as "ask anything about our documents", and nobody could say in advance what a good answer would look like. A system with no agreed target cannot miss it, and cannot hit it either. Pilots tied to a specific question with a named owner and a reference answer fail more informatively than open ones.
The second is that no test was agreed before the pilot began. Success was judged by demonstrations, and the most persuasive question in a demonstration is the one the sponsor happens to ask. Without an evaluation set written down in advance, every post-mortem is an argument about anecdotes, and the loudest anecdote wins.
Where this advice is too tidy
The separation is cleaner on paper than in practice. Some corpus problems are engine problems in disguise: a good engine can surface a contradiction between two documents so that a person resolves it, where a poor one hides it. And the order of repair matters. Improving the engine first, on a bad corpus, produces fluent wrong answers, which are worse than clumsy wrong ones because they are easier to believe. Where both are broken, fix the corpus problems that cause the most wrong answers first, and re-run the test.
The diagnostic also costs effort. Someone has to find the answers by hand for a few dozen questions. That effort is small beside another quarter of unexplained disappointment, but it is not zero, and it needs a person who knows the material.
Before the next pilot
Write the evaluation set before configuring anything. Use real questions from the people who will use the system, with reference answers and the location of each answer in the corpus, plus a handful of questions that have no answer. Agree in advance what proportion has to be right, and what the system should do when it cannot tell. Run the diagnostic two weeks in, not six months in.
In our own discovery work we build the evaluation set from the client's questions and records before anything is configured, because nothing else lets a result be explained afterwards. Whether you work with us or not, the principle holds. A pilot that cannot say why it failed has not produced anything you can use, and one that can say so has already paid for part of the next attempt.
Talk to us
