A pilot shows that a model can answer. Production requires that it answers correctly, for the right people, at a cost someone has agreed to pay.
Most stalled AI programmes did not fail on model quality. They stalled because the pilot was run under conditions that production will not offer. The demonstration was genuine, and it answered a different question from the one the business later needed answered.
Compare the two settings.
| Pilot condition | Production condition |
|---|---|
| A curated extract of clean data | Live systems with duplicates, gaps and late updates |
| A dozen prepared questions | Thousands of unprepared ones, including ambiguous and out-of-scope |
| A few friendly users | Many users with different entitlements |
| One administrator with full access | Per-user policy on every query |
| Cost ignored | A budget with an owner |
| Accuracy judged by impression | Accuracy measured against agreed answers |
| A sponsor who likes the result | An accountable owner for the outcome |
Each row is a place where the pilot hid a problem. The curated extract concealed the join errors between systems. The prepared questions concealed how the system behaves when it does not know. The single administrator concealed the entitlement problem entirely, and entitlement is often the issue that stops a security review and ends the programme.
There is a sequence behind this. Making data answerable comes first. Making AI usable on that data comes second. Letting agents do the routine work comes third. Programmes that attempt the third stage before the first is complete tend to stall, because an agent acting on data nobody trusts only produces wrong actions faster.
Production also introduces work that a pilot never meets. Someone must monitor answer quality over time as data and models change. Someone must handle the questions that fail. Someone must pay for the compute. These are operating tasks and they continue indefinitely, so a pilot that has no plan for them has only postponed the cost.
You can design a pilot to learn about production instead of to impress. Use real data from live systems, even if the scope is narrow. Include users with restricted access. Write the success criterion down before the work starts, in terms of the business question and not of the technology. Agree in advance that a verdict that the work should not proceed is an acceptable result. A pilot that cannot return that verdict is a demonstration.
Finally, ask who will own the result. A pilot with a sponsor but no owner usually ends when the sponsor's attention moves on.
In practice: Before approving the next pilot, go through the table above and mark each row as it will apply to the pilot. Every row that reads as a pilot condition is a risk you are choosing to carry into the production decision. Fix the cheap ones now, starting with live data and restricted users.
Next: B2, A citation shows where an answer came from, and evidence shows that it is supported.
Talk to us
