EigenForge AI Labs Talk to us
Sovereign AI — track artwork

Learn · Track D

Sovereign AI

Your context is the asset. What sovereignty means beyond on premises, and how to keep it home.

Track D · part 1 of 5

D1. Sovereignty is a property of where data, inference and evidence live, not of where the server sits

A deployment is sovereign only if content, model inference and the audit trail all stay inside your boundary.

"We can deploy on premises" is the standard vendor answer to a sovereignty question. It is true for most vendors and tells you little. A product can run in your data centre and still send content out to be embedded, call a hosted model for every answer and write its audit trail to the vendor's tenancy.

Four places data can escape

Trace each of these separately.

  • Indexing. Many AI systems copy your content into a vector store, and some embed it using a hosted service. Ask where chunking and embedding happen and whether the copy lives inside your boundary.
  • Inference. The model that answers the question may run in a region you cannot name. Ask where inference runs and whether it can run against a model you host.
  • Control plane. Policy, entitlement and licensing logic may sit in the vendor's cloud. If it is unreachable, does the product still run?
  • Evidence. Logs, traces and audit records may be written to the vendor's tenancy. You can see your data, but you cannot prove what happened to it.

A fifth route is quieter: telemetry, licence checks and periodic syncs. A deployment described as air-gapped but making a scheduled outbound call is not air-gapped.

The product test

The second question is whether the on-premises version is the same product as the hosted one. Some are diminished editions: a cloud product ported to run locally, trailing the hosted release, missing features and dependent on managed services that have no local equivalent. A sovereign deployment of a different and lesser product is a compromise, and you should be told so before the contract.

A practical test is to list the capabilities you saw in the demonstration and ask which of them work with the network cable removed.

Sovereignty is several rules, not one

The rule that applies to you may be a regulation, a data residency requirement, a classification level or a contract clause with your own customers. Each constrains different things. A residency rule may care about storage location only. A classification regime may restrict who can administer the system. A customer contract may prohibit the use of customer content to train or tune a model. State the rule precisely before you evaluate vendors, because it decides which of the four escape routes matter.

Read access and blast radius

A further design question is how the AI reaches your data. A system that reads records in place, read-only, with no migration and no write-back, has a small footprint and is easier to approve. A system that requires you to move data into its store creates a second copy that you must secure, reconcile and eventually delete.

Apply the switch-off test. If you disconnect the system, does everything else carry on as before? If not, it has become a dependency, not an added layer.

In practice: Take the shortlist from your current evaluation and ask each vendor, in writing, the eight questions that follow from this module: where content is embedded, where inference runs, where policy is evaluated, where the audit trail lives, what stops working when disconnected, what depends on a specific cloud provider, whether the release is identical to the hosted one, and what the switch-off test shows. Compare answers, not brochures.

Next: D2, Five topologies and what each one really costs.

Put it to work

The common mistake. Accepting "we can deploy on premises" as the sovereignty answer. The server sits in your building, but content is embedded by a hosted service, inference calls an external model, and the audit trail is written to the vendor's tenancy.

How to check. Put the eight written questions to each vendor. Then run the cable test in a trial environment: list the capabilities from the demonstration, disconnect the network, and note which stop working. Watch outbound traffic for scheduled calls.

What good looks like. Indexing, inference, policy evaluation and evidence all sit inside your boundary, the release matches the hosted one, and disconnecting the system leaves everything else running as before.

YOUR SOVEREIGN BOUNDARYSYSTEMS OF RECORD — READ IN PLACEERPCore platformMainframeCRMFiles & streamsGOVERNED LAYERDefinitions · quality · lineage · entitlement evaluated at query timeONTOLOGY GROUNDINGAnswers generated from a model of how yourtables relate — no vector copy, nothing to drift.Nothing is chunked and shipped out.LOCAL MODEL INFERENCEOpen-weights models on your own hardware.Model choice is a configuration, not an architecture.No inference call leaves the boundary.OUTSIDE THE BOUNDARYNothing. No embedding service, no hosted control plane, no telemetry channel, no licence call-home.
Exhibit 01Everything inside the boundary. Nothing outside it.

Track D · part 2 of 5

D2. Each deployment topology trades control against operating effort, and the trade is predictable

On-premises, private cloud, public cloud, air-gapped and federated deployments differ mainly in who carries the operational load and how much change you can absorb.

A topology is a choice about where the system runs and who operates it. The same software can often run in all five, if it avoids dependence on any one provider's managed services. The costs differ, and they are mostly operational.

The five topologies

TopologyControlMain cost to plan for
On premisesFull, inside your estateHardware, accelerators, patching, staff
Private cloudHigh, your tenancy and keysKey management, tenancy design, provider dependence
Public cloudModerate, on standard servicesEgress, residency configuration, spend control
Air-gappedHighestTransfer of updates and models, no remote support
FederatedPer jurisdictionCoordination across sites, consistent policy

On premises. You supply the hardware, including accelerators for model inference, and the people who run them. Capacity planning is yours. This is the usual choice where data cannot leave the building.

Private cloud. The system runs in a tenancy you control, with keys you hold. You gain elasticity without sharing infrastructure. You still depend on the provider for availability and must verify that the provider's staff cannot read your keys.

Public cloud. Any provider will do if the system runs on standard container orchestration and avoids provider-specific databases, queues and identity services. The costs that surprise teams are data egress, accidental use of the wrong region and unmetered model spend.

Air-gapped. Nothing connects to the outside. Inference must run on local models. Software updates, model weights and security patches arrive through a controlled transfer, usually on physical media after scanning. Plan the update cycle early, since it decides how current your system can stay. Remote support is not available, so documentation and local skills matter more.

Federated. Each jurisdiction runs its own deployment, and a group-level view is assembled by querying them. D4 covers this in detail.

Hidden costs that apply to all five

  • Accelerators. Model size sets hardware need. Lead times for accelerators can be long, so plan procurement alongside design.
  • Evaluation. Each topology changes the available models, which means re-running your evaluation, not assuming results carry over.
  • Upgrade discipline. Every site you operate is a site you must keep current. Divergent versions are the main source of quiet failure.
  • People. Someone must own backups, monitoring and incident response in each environment.

Choosing between them

Start from the constraint, not the preference. If a rule forbids content leaving a boundary, the topology follows from where the boundary is. If there is no such rule, choose the least operational effort that meets your risk appetite, which is often public or private cloud.

Resist the idea that one topology is permanent. Requirements change with new regulation, new customers and new markets. A product that runs identically across topologies lets you move. A product built for one makes the move a re-implementation.

In practice: For each system you plan to run, write down the binding constraint, the topology it implies and the three costs from the list above that you have not yet budgeted. Ask vendors whether each topology is the same codebase and the same release, and how updates reach an air-gapped site.

Next: D3, Choosing open-weights models without a leaderboard.

Put it to work

The common mistake. Choosing a topology by preference and budgeting only for the software. Accelerator lead times, the transfer of updates to an air-gapped site, and the staff for backups and incident response surface after the contract is signed.

How to check. For one system, write down the binding constraint, the topology it implies and the three hidden costs you have not budgeted. Ask the vendor how an update reaches an air-gapped site and whether each topology runs the same release.

What good looks like. The topology follows from a stated constraint. Accelerators are ordered alongside design, each site has a named operator, versions are kept aligned, and moving between topologies is a configuration exercise and not a rebuild.

01On premisesInside your data centre.Nothing leaves the estate.Defence · public sector ·regulated manufacturing02Private cloudYour tenancy, your keys.Financial services ·healthcare03Public cloudAny provider, standardKubernetes.Cloud-first enterprises04Air-gappedDisconnected, with localmodel inference.Classified · criticalnational infrastructure05FederatedPer jurisdiction, one viewacross them.Multi-country groups underresidency rulesFive topologies from one codebase — not a cut-down edition six months behind.
Exhibit 02Five topologies from one codebase.

Track D · part 3 of 5

D3. Choose the smallest model that passes your own evaluation, and read its licence first

Public leaderboards measure what the benchmark authors chose to measure, and your workload is rarely that.

In a sovereign deployment you run inference on models you hold. Open-weights models make that possible. They also present a long and fast-changing list of options, and published rankings are a poor guide to which one will answer your questions well.

This module avoids naming specific models or versions because they date within months. The method does not.

Four shapes of model

Most deployments use two or three of these, not one.

ShapeTypical roleWhat to test
General instruction-tuned, mid-sizeGrounded question answering and summarisationInstruction-following consistency under your own prompts
Long-contextDocument extraction, contracts, case filesPerformance at the far end of the context window, which is often weaker than the stated figure
Larger, reasoning-capableAgent planning and multi-step tool useReliability of tool calls and plan quality
Small or code-specialisedClassification, routing, coding help, on-device workWhether a simple rule would do the job more cheaply

Selection criteria, in order

  1. Your evaluation set. Build it from your own questions and records. Include questions the system should refuse, and cases where the right answer is "the evidence does not support one". Nothing else should decide the choice.
  2. Latency and throughput. A model that scores slightly better and runs several times slower is often the wrong model for the workflow.
  3. Hardware. Use what you have or are prepared to buy. If a smaller model on existing hardware meets the bar, choose it.
  4. Licence. Have someone who will need to defend it read the terms. Some open-weights licences restrict commercial use, field of use, scale or redistribution in ways a regulated organisation cannot accept. Open weights and open source are not the same thing.
  5. Operating cost. Measure the cost of each answered question, including retries and failed attempts, not the price per token.

Building the evaluation set

Collect real questions from the people who will use the system. Pair each with the record or document that answers it, and a reference answer written by a subject expert. Add adversarial cases: questions that tempt the model to invent a source, questions that need two documents, and questions the user is not entitled to have answered. Keep a portion unseen, so tuning does not overfit.

Score for correctness, grounding in the cited source, correct refusals and format compliance. Re-run the set whenever the model, the prompt or the data changes.

Keep the choice reversible

Design so that the model is a configuration. Grounding, policy, entitlement and evidence should live in the platform around the model. Then replacing the model means re-running your evaluation, not rebuilding the system. Record what you tested, what you rejected and why, so the next review starts from evidence.

In practice: Assemble a first evaluation set of a modest size from real questions before you shortlist any model. Test two candidate shapes against it on your own hardware, and record answer quality, latency and cost per answered question. Ask your legal function to read the licence of the leading candidate before the pilot starts.

Next: D4, Federation: one view without one location.

Put it to work

The common mistake. Picking the model at the top of a public ranking and reading the licence after the pilot. Legal then finds a restriction on commercial use or field of use, and the evaluation never tested refusals or long documents.

How to check. Ask your legal function to read the licence of your leading candidate this week. In parallel, run ten of your own real questions, including two that should be refused, through two model shapes on your own hardware, and compare answers and latency.

What good looks like. The model is a configuration setting. The smallest one that passes your evaluation set on your hardware is in use, its licence is cleared, and a record of what was rejected and why is ready for the next review.

WORKLOADWHERE WE USUALLY STARTWHYGrounded question answeringMid-size instruction-tuned openweightsAccuracy on your ontology matters more than rawscale; the retrieval does the heavy lifting.Document extractionSmall to mid-size, long-context openweightsContext length and structure-following beat reasoningdepth.Agent planning and tool useLarger open weights, or a hostedmodel where policy allowsTool-calling reliability is the constraint; test itrather than assume it.Coding assistanceCode-specialised open weights, runlocallyLatency and privacy dominate; a smaller specialisedmodel beats a larger general one.Classification and routingSmall open weights, or no model atallMost routing is a rule. Reach for a model only wherea rule genuinely fails.HOW WE CHOOSE✓Smallest model that passes your evaluation, not thelargest you can run✓Permissive licence you have read, and weights you canhold✓Measured on your data, not on a public leaderboard
Exhibit 03Where we start, by workload.

Track D · part 4 of 5

D4. Federation answers group questions by asking each site, not by gathering the data

A federated design keeps each jurisdiction's data where its rules require, and builds the group view from answers rather than from copies.

The hard sovereign requirement is rarely "keep it in the country". It is "keep it in the country and still let the group see the whole picture". Centralising data into a group warehouse meets the second and breaks the first. Leaving each country alone meets the first and leaves the group blind.

How federation works

Each jurisdiction runs its own deployment. It holds its own data, applies its own entitlements and follows its own rules. A group-level question goes to a coordinating layer, which breaks it into sub-questions, sends each to the relevant deployments, receives the answers and combines them.

The group layer never holds the underlying records. It holds the question, the returned results and enough metadata to route and reconcile them.

Four design decisions

  • What leaves a jurisdiction. Decide, per jurisdiction, whether the answer may be a record, an aggregate, a count or nothing. Aggregation thresholds prevent small counts from identifying individuals.
  • Where redaction happens. Redaction must happen inside the jurisdiction before the answer leaves. If the group layer redacts, the sensitive data has already crossed the border.
  • Whose entitlements apply. A group analyst has rights at group level that do not necessarily extend to each country's records. Evaluate the user's entitlement in each jurisdiction, so the answer reflects what that user may see there.
  • Which definitions are shared. A federated answer is only meaningful if "active customer" or "overdue invoice" means the same in each country. A shared set of business definitions, versioned and owned centrally, is more important than any technology choice.

What federation cannot do

Federation is not free. A query across sites is only as fast as the slowest site. If one is unavailable, the group view is incomplete, and the system should say so instead of silently returning a partial figure. Analysis that needs row-level data from every site, such as training a model on pooled records, does not fit the pattern and may need techniques designed for that purpose or may not be permitted at all.

Consistency is a further limit. If sites hold data at different freshness, a group figure combines different moments. The answer should state the as-of time for each contributing site.

Governance across sites

Each site keeps its own audit trail, in its own deployment. The group layer records which questions were asked, which sites answered and what was released, and these records are sealed in the way C4 describes. A regulator in any country can then review the local record without access to the others.

Version discipline matters as much as architecture. A group that runs different releases in different countries will find that the same question returns inconsistent answers, and the cause will take time to find.

In practice: Map your jurisdictions in a table with three columns: what may leave, the form it may take, and who may ask. Agree the shared definitions for your ten most-used group metrics before choosing technology. Then test one group question end to end, including the case where one site is unavailable.

Next: D5. Your context is the asset — sovereignty is keeping it home.

Put it to work

The common mistake. Letting the group layer do the redaction, so the sensitive record has already crossed the border by the time a field is removed. A second version is comparing figures across countries that define "active customer" differently.

How to check. Pick one group question and trace what leaves each jurisdiction, in what form, and where redaction occurs. Then collect each country's definition of your ten most-used metrics and compare them. Disconnect one site and see what the group view reports.

What good looks like. Per-jurisdiction rules state what may leave and in what form, redaction happens inside, entitlements are evaluated locally, and definitions are shared and versioned. A missing site is reported as missing, with an as-of time for each contributor.

Jurisdiction Aown deploymentown data, own rulesJurisdiction Bown deploymentown data, own rulesJurisdiction Cown deploymentown data, own rulesOne federated viewasks each deployment — it never gathers the dataRedaction is applied per jurisdiction on the way out, so a group answer never contains something a jurisdiction would not release.
Exhibit 04One view, without one location.

Track D · part 5 of 5

D5. Your context is the asset — sovereignty is keeping it home

Sovereignty is usually sold as where the server sits. The sharper question is where your context goes: the records, relationships, definitions and history that make answers about your organisation possible.

A record is a fact. Context is what makes the fact answer a question: which customer it belongs to, how that customer relates to three other entities, which definition of revenue your board agreed, what the person asking is allowed to see. Context is the accumulated, organised memory of how your organisation works, and it is the one part of an AI system that is not available for purchase.

When an AI answers a question about your organisation, the answer's quality is mostly the context's quality. Models are a market; context is an asset. Whoever holds your context holds the thing that makes every future answer possible — and if it lives inside a vendor's system, assembled in their format, improving their product, you have rebuilt the lock-in of the last decade one layer up the stack.

Context rarely leaves dramatically. It leaks politely, a request at a time. An agent that sends the whole ticket, customer record and account history to an external API with every question is exporting context. A retrieval index on someone else's infrastructure is exported context. A copilot that learns your definitions and thresholds inside a vendor's cloud is exporting the most valuable layer — discovered, usually, on the day you try to leave.

The fix is to draw the boundary correctly and be precise about what crosses it. Records never cross: questions resolve against an ontology over live records inside your perimeter. Inference is a choice: a model running inside your boundary means nothing crosses; an external model means a contracted, minimal crossing — and both should be configurations of the same codebase, so sovereignty is not a diminished edition. Evidence stays home: the record of what was asked, answered and decided is append-only and yours. Modules D1 and D2 cover the deployment topologies; this module is about the asset they exist to protect.

In practice: Ask each AI vendor you use one question — if we leave, what do we take? Raw data and a farewell email means the context was never yours. The ontology, definitions, evaluation sets, decision records and the means to run the same queries elsewhere means it was.

Next: E1. A question the board can approve names a decision, an owner, a measure and a date.

Put it to work

The common mistake. Certifying the data centre region and calling the system sovereign, while a copilot assembles your definitions, thresholds and naming conventions inside a vendor's cloud — the layer you will miss most on the day you try to leave.

How to check. Ask each vendor the exit question: if we leave, what do we take? Then trace one real question end to end and list what crosses the boundary with it — records, history, definitions, or only a task-shaped question stripped of all three.

What good looks like. Records never cross, inference location is a configuration of the same codebase, and the evidence store is append-only and yours. The exit answer includes the ontology, definitions, evaluation sets and decision records.

FIG. 1 THE CONTEXT BOUNDARY — WHAT A MODEL PROVIDER EVER SEESINSIDE YOUR PERIMETER — YOUR INFRASTRUCTURE, YOUR RULESSYSTEMS OF RECORDread in place, never copied outONTOLOGY & DEFINITIONSagreed once, enforced everywhereINFERENCEa model you choose, running hereEVIDENCE STOREappend-only, hash-chained, yoursTHE BOUNDARYA TASK-SHAPED QUESTION, STRIPPED OF RECORDSA REASONED RESULT, CHECKED AGAINST YOUR EVIDENCE HEREOUTSIDEMODEL PROVIDERsees a question, never your records;retains nothing by contract and byarchitectureOr nothing crosses at all: the modelruns inside too — an ordinaryconfiguration, not a diminishededition.
Exhibit 05What crosses the boundary, what does not, and what stays home.

Something here you disagree with?

These are written to be argued with. If a part of this is wrong, or missing the case you care about, tell us and we will fix it.

hello@eigenforgelabs.ai

Send opens your email client with the note already addressed to us — nothing is stored on this site, and the message goes from your own mailbox, so our reply lands in yours.