An evaluation runs for three months. A favourite emerges. Then someone from security reads the architecture document properly and finds that content has to leave the building to be embedded, or that inference happens in a region nobody on the vendor's side can name, or that the audit trail lives in the vendor's tenancy. The conversation turns to exceptions, mitigations and legal opinions, which is an expensive way to learn that the shortlist was wrong.
It happens often enough that it looks like a pattern, and it begins with a question that sounds sensible and tells you almost nothing: can this be deployed on premises? Every vendor says yes. The word costs nothing, because it describes where some software runs and says nothing about what the software does, what it depends on, or where the data goes while it does it.
This essay is about replacing that question with better ones. They are, first, whether the on-premises version is the same product as the hosted one, and second, whether the AI still works when nothing can leave.
Who this is for, and who it is not
Not every organisation needs sovereign deployment, and that deserves saying before the argument begins.
Where the constraint is a statute, the route may be a contract. Singapore's Personal Data Protection Act restricts transfers of personal data outside the country, but the transfer limitation obligation in section 26 allows a transfer if the organisation takes appropriate steps and the recipient is bound by legally enforceable obligations, including contractual ones, to protect the data to a standard comparable to the Act's. The PDPC's guidance.pdf) sets out the conditions and the alternative pathways. For many organisations a hosted service with the right contract, in the right region, is a legitimate answer and often the cheaper one.
The rest of this essay is for those whose constraint is firmer than that: a classification level, a contract with their own customers that forbids it, a regulator's expectation, or a board that will not have it. If that is you, the market's default answer is already largely closed, and the remaining question is how to tell which of the offered alternatives is real.
Five ways sovereign AI disappoints
In our experience the failures fall into five patterns, and a vendor can show any one of them while sounding entirely credible.
The first is the diminished edition. The vendor has a cloud product, and the on-premises version is a port of it. It trails the hosted release by a cycle or three, some capabilities are missing, and the roadmap belongs to the cloud product. You are buying yesterday's software at tomorrow's price, and the gap widens, because the engineering effort follows the customers who pay for the hosted version.
The second is that the AI is still somewhere else. The application runs inside your boundary, and the intelligence does not. Content is split into chunks and sent elsewhere to be embedded, or the question and its context go to a hosted model for the answer. The deployment diagram is sovereign and the data flow is not. This is the commonest pattern because it is the easiest to hide: the part that does the thinking is a single API call that rarely appears on the architecture slide.
The third is the managed-service dependency. The product assumes a particular cloud provider's managed database, queue, identity service or model endpoint. Lift it out of that provider and half of it fails to start. Such products can be installed on premises in the sense that the installer runs, and still not work there.
The fourth is that governance follows the data to the vendor. The audit trail, the policy engine and the entitlement model live in the vendor's control plane. You can see your data, but you cannot prove what happened to it, because the record of what happened is held by someone else. Of all five this does the most damage over time, because it surfaces when something goes wrong and a regulator asks for the record.
The fifth is that nothing is disconnected in fact. "Air-gapped" turns out to mean a periodic sync, a licence check, a telemetry channel somebody assumed was harmless, or a safety filter that calls a cloud service. The AI sovereignty maturity model from Traefik makes a point that applies well beyond its own product: self-hosted but connected is a different level from self-hosted and autonomous, and the useful test is whether everything keeps working if you pull the network cable. Its more general principle is that your overall level of sovereignty is set by your weakest dimension. A single external dependency caps the whole architecture.
Eight questions
The questions that follow are the ones that separate the five patterns. They are best put in writing, to every vendor on the list, and compared.
Start with identity. Is the on-premises version the same product and the same release as the hosted one? This catches the diminished edition. If the answer is "mostly", ask what is missing, and ask for it in writing. A related question catches the managed-service dependency: what does the system depend on from a specific cloud provider? A good answer is a short list that you can supply yourself, such as a standard container platform. A poor answer is a list of that provider's product names.
Then ask where the AI is. Where is content embedded, and does anything leave the boundary to be indexed? And where does inference run, and can it run against a model we host ourselves? These catch the second pattern. They are better asked as requests to see the network path than as questions about it, since a design document describes intentions and a packet capture describes behaviour. The aim is to be told which components make outbound calls, to where, and carrying what.
Next comes governance. Does the audit trail live in our deployment or in yours? Can policy be evaluated per user at the moment of the query, or only at the front door? Can you reconstruct a decision from the data as it stood on the day it was made? Each of these is about whether control stays with the data. The second matters more than it looks. Policy applied only at login means that two people with different entitlements may receive the same answer, or that someone maintains separate extracts for different audiences, which then drift. The third is the one an auditor will eventually ask, and it requires that the history of the data, the model and the prompt be kept in a form you can replay.
Finally, ask what stops working when the network is cut. In an air-gapped deployment, which capabilities stop? The right answer is a specific list, even if a short one. The wrong answer is "nothing", delivered without a demonstration. Whatever the answer, test it.
Taken together, the eight reduce to two. Is it the same product, and does the AI still work when nothing can leave? The other six are ways of checking those.
What doing this properly costs
The questions are easy to ask, and the consequences are not free, so leaving them out would mislead you.
A model you host yourself will, for some tasks, be less capable than the largest hosted models. For a lot of enterprise work, such as grounded question answering over your own documents, summarising, classification, extraction and routing, a mid-sized open-weight model that passes your own evaluation is often enough, and a model that is slightly better on a public leaderboard is not the right test. But for open-ended reasoning and long multi-step planning the gap can matter, and the right response is to measure it on your own questions, not to assume it away. The licence deserves reading as well: some open-weight models carry conditions that a regulated organisation cannot accept.
Hardware has to be bought and operated. Updates arrive on your schedule, which is a security advantage and an operational burden. In a fully disconnected environment, every model update, patch and licence renewal is a physical process. These are real costs, and a vendor who pretends otherwise is describing the first pattern.
There is also a case for a mixed arrangement. Not every workload needs the same boundary. A sensible approach classifies workloads by what they touch: material that can never leave stays inside, and the rest may use hosted services under a suitable contract. That makes the boundary a decision per workload, not a single ideological position, and it makes the eight questions matter most for the workloads that sit inside it.
How to run the diligence
Ask in writing, and ask everyone on the shortlist the same eight questions, so the answers can be compared. Ask for the egress list: every outbound connection the product makes, in normal operation and during start-up, upgrade and licence validation. Then test it. Block outbound traffic in a trial environment and see what fails. A short proof of concept in the topology you will use, including a disconnected one if that is your requirement, with a success criterion written down beforehand, will settle more than a long architecture review. It will also tell you something about the vendor: whether they welcome the test or negotiate around it.
Read the answers for specificity. "Fully supported" and "enterprise ready" carry no information. "These four components make outbound calls, to these destinations, for these reasons, and each can be disabled" does. Any vendor who can answer that way has already done the work.
Closing
The on-premises question persists because it is easy for both sides. The buyer can tick a box, the seller can answer yes, and the real discussion is deferred until the point at which deferring it is most expensive. The better questions take longer to ask and sometimes produce an uncomfortable silence, which is the information you wanted.
We are asked these eight questions often, and we are glad to answer them on the record, in writing. We would encourage you to put them to us, and to everyone else on the list, and to compare the answers. If a vendor's answers are specific, testable and the same on a Tuesday as they were on a Friday, you can proceed. If they are not, you have saved yourself the three months.
Talk to us
