A deployment is sovereign only if content, model inference and the audit trail all stay inside your boundary.
"We can deploy on premises" is the standard vendor answer to a sovereignty question. It is true for most vendors and tells you little. A product can run in your data centre and still send content out to be embedded, call a hosted model for every answer and write its audit trail to the vendor's tenancy.
Four places data can escape
Trace each of these separately.
- Indexing. Many AI systems copy your content into a vector store, and some embed it using a hosted service. Ask where chunking and embedding happen and whether the copy lives inside your boundary.
- Inference. The model that answers the question may run in a region you cannot name. Ask where inference runs and whether it can run against a model you host.
- Control plane. Policy, entitlement and licensing logic may sit in the vendor's cloud. If it is unreachable, does the product still run?
- Evidence. Logs, traces and audit records may be written to the vendor's tenancy. You can see your data, but you cannot prove what happened to it.
A fifth route is quieter: telemetry, licence checks and periodic syncs. A deployment described as air-gapped but making a scheduled outbound call is not air-gapped.
The product test
The second question is whether the on-premises version is the same product as the hosted one. Some are diminished editions: a cloud product ported to run locally, trailing the hosted release, missing features and dependent on managed services that have no local equivalent. A sovereign deployment of a different and lesser product is a compromise, and you should be told so before the contract.
A practical test is to list the capabilities you saw in the demonstration and ask which of them work with the network cable removed.
Sovereignty is several rules, not one
The rule that applies to you may be a regulation, a data residency requirement, a classification level or a contract clause with your own customers. Each constrains different things. A residency rule may care about storage location only. A classification regime may restrict who can administer the system. A customer contract may prohibit the use of customer content to train or tune a model. State the rule precisely before you evaluate vendors, because it decides which of the four escape routes matter.
Read access and blast radius
A further design question is how the AI reaches your data. A system that reads records in place, read-only, with no migration and no write-back, has a small footprint and is easier to approve. A system that requires you to move data into its store creates a second copy that you must secure, reconcile and eventually delete.
Apply the switch-off test. If you disconnect the system, does everything else carry on as before? If not, it has become a dependency, not an added layer.
In practice: Take the shortlist from your current evaluation and ask each vendor, in writing, the eight questions that follow from this module: where content is embedded, where inference runs, where policy is evaluated, where the audit trail lives, what stops working when disconnected, what depends on a specific cloud provider, whether the release is identical to the hosted one, and what the switch-off test shows. Compare answers, not brochures.
Next: D2, Five topologies and what each one really costs.
Talk to us
