EigenForge AI Labs Talk to us
Agents and governance — track artwork

Learn · Track C

Agents and governance

Identity, authority, harnesses and evidence for principals who are not people.

Track C · part 1 of 6

C1. A system is agent-native when the agent has an identity, a limit and a record of its own

The difference between agent-added and agent-native is whether the underlying system knows the agent exists.

Most enterprise software now has an assistant. A chat panel sits beside an application that was designed years ago for people with logins. The assistant can read what is on screen, summarise it and suggest a next step. It is an agent-added design. The agent stands outside the system and looks in.

An agent-native system treats the agent as a participant. The system recognises it as a distinct identity. Someone has granted it a defined scope of authority. It has a budget. There is a route for approval when it reaches the edge of that scope. When it acts, the record shows the agent by name, not a person's session or a shared service account.

The difference sounds like a matter of degree. It is closer to a difference of kind. An agent-added system can advise. An agent-native system can own a task, because there is something to hold responsible.

Four properties the substrate must have

The test is whether four things exist in the system itself, not in a prompt or a policy document.

  • Identity. The system can tell which agent acted, and distinguish it from every person and every other agent.
  • Authority. There is a defined answer to what this agent may do, stored where the system consults it at the moment of action.
  • Boundary. A mechanism stops the agent at the edge of its authority. An instruction in a prompt asking it to behave is not a mechanism.
  • Record. What the agent did is written somewhere the agent cannot alter.

Software built for human users usually has none of these for a non-human actor. Its identity model assumes a person. Its permission model assumes a session that a person opened. Its logs record what the application did, not who directed it. You can attach an assistant to such a system. You cannot attach the concept of a principal that can be held to account.

Why retrofitting fails

The common retrofit is a service account with broad permissions, an API key and a logging wrapper. Each piece looks reasonable. Together they produce an actor that nobody owns, whose permissions were set once by an engineer who has since moved teams, and whose actions appear in the log under a name that means nothing to an auditor.

A second failure is enforcement at the wrong point. If the boundary is checked only when a user opens the application, then scheduled jobs and background processes pass around it. Agents run mostly in the background. The check has to sit on the decision path, where the action is taken.

How to evaluate any product

Ask the vendor to show, not describe, the following. Where is the agent's identity created, and who can see it? Where is its authority stored, and can a business owner change it without a deployment? What happens when it attempts something outside its authority? Can it write to its own record?

If the answers involve a prompt, a wrapper or a convention, the product is agent-added. That may be acceptable for drafting and search. It is a poor foundation for work with financial or operational consequence.

In practice: Pick one process you would like an agent to run. Write down which of the four properties your current system provides for it and which it does not. The gaps tell you whether to extend the system or keep the agent in an advisory role.

Next: C2, Identity, authority and budget for principals who are not people.

Put it to work

The common mistake. Calling a system agent-native because the assistant has a login. Teams give it a service account with broad permissions and a logging wrapper, so its actions appear under a name that means nothing to an auditor and nobody owns the permissions.

How to check. For one process you want an agent to run, ask the vendor to show, not describe, four things: where the agent's identity is created, where its authority is stored, what happens at the boundary, and whether it can write to its own record.

What good looks like. The agent has its own identity, authority held where the system consults it at the moment of action, a mechanism that halts it at the boundary, and a record it cannot alter.

Channelsvoice, chat, web, analyst co-working, offline field clientAgent runtime & orchestrationapproval and interrupt as primitives, long-running workflowsPolicy engine & approvalsaction gates on the decision path, approval routingEvidence & attestationsealed record per action, artefact hashing, model provenanceMemory, graph & datagraph and vector memory, governed lake, hybrid search, temporalityConnectors & sandboxed executionopen protocols, container and micro-VM sandboxesACROSS EVERY LAYERIdentity and entitlement forpeople and agents alikeAny-model routing, withself-hostingOn premises, at the edge, orair-gappedResidency and redaction byjurisdictionEncrypted mesh between sitesPerimeter-native: the engine runs inside your boundary and the audit is signed at the edge.
Exhibit 01The layers an agentic system needs before it is allowed to act.

Track C · part 2 of 6

C2. An agent needs an owner, a scope and a budget before it needs a model

Treat an agent as a principal: someone is named as responsible, the scope is written down, and spending stops at a limit.

A principal is any actor the system recognises and holds to its rules. People have always been principals. Agents now need to be, and the controls that organisations apply to staff translate well if you follow them one at a time.

Identity: one agent, one identity

Each agent gets its own identity in the same directory or identity layer that people use. It should not borrow a person's credentials, and it should not share a service account with other agents. Shared identities make it impossible to say which agent did what, and they make it impossible to withdraw access from one without affecting the rest.

Every agent identity needs a named human owner. The owner is the person who answers for the agent's behaviour, grants it more authority and withdraws it. An agent whose owner has left the company should be suspended automatically until a new owner is named.

Authority: scope by action, not by system

Authority is easier to reason about when it is expressed as actions on objects. "May read invoices for this entity. May draft a payment proposal. May not release a payment." This is more precise than "has access to the finance system", and it maps to the action gates covered in C3.

Inherit the entitlements that already govern people where you can. If a user cannot see a record, an agent acting for that user cannot see it either. Where an agent acts on its own account, give it a scope that is deliberately narrower than that of the person who supervises it.

Grant authority in steps. A new agent starts by advising. After a period in which its output has been reviewed, the owner may allow it to prepare actions for approval. Autonomy is extended only for action classes where the evidence supports it.

Budget: a limit that refuses work

Agents consume money whenever they call a model, a tool or an external service. A loop that retries a failing task can spend a month's allowance in an afternoon. A dashboard that reports the overrun the next morning does not help.

A budget should be a control, not a report. It is attached to the agent, the team or the use case, and when it is reached the work is declined, with the refusal recorded. Someone then decides whether to raise the limit. The decision becomes a record in its own right.

A short checklist

QuestionAcceptable answer
Who owns this agent?A named person, listed in the inventory
What can it do?A written list of actions on defined objects
What stops it?A policy check on the decision path
What does it cost?A limit per agent, with refusal at the limit
Who can revoke it?The owner and a central administrator, immediately

In practice: Take your current list of agents, including pilots and scripts that call a model, and add four columns: owner, scope, budget and last review date. Any row with an empty cell is a risk that a regulator or auditor will find first. EigenForge is building this model into Mnemos, which is in development, and the same questions apply to any platform you assess.

Next: C3, Action gates: advise, approve, auto, and how to decide which.

Put it to work

The common mistake. Starting with the model and leaving ownership for later. Pilot agents and scripts that call a model accumulate with no named owner, no written scope and no spending limit, and the engineer who set them up changes team.

How to check. Build an inventory this week, including pilots and scripts. Add columns for owner, scope, budget and last review date. Then ask each owner to say what the agent may not do. Every empty cell is a finding.

What good looks like. Every agent has its own identity, a named owner, a written action list, a budget that refuses work at its limit, and a review date. An agent whose owner leaves is suspended until a new one is named.

Track C · part 3 of 6

C3. Decide the gate for each action class before anything is built

Advise, approve and auto are three levels of autonomy, and the right one depends on the consequence and the reversibility of the action.

Autonomy should be granted by action class, not by agent. The same agent may send a reminder on its own, prepare a refund for approval, and only advise on a customer's eligibility for a benefit. Each of those is a different class with a different consequence.

The three gates

Auto. The agent acts and the action is recorded. This suits routine, reversible actions within set limits. Examples are sending a reminder, classifying a document and moving a file to a folder. If the action is wrong, the cost of undoing it is low.

Approve. The agent prepares the action and assembles the evidence for it. A named person reviews and approves before anything happens. This suits most work with financial or operational consequence, such as a payment proposal or a change to a supplier record.

Advise. The agent informs, and the person decides. This suits any case where an individual's entitlement, eligibility or liability is at stake. The agent may gather facts, cite rules and set out options. It does not grant, deny or alter.

Four questions that place an action in a gate

  1. Can it be reversed, and at what cost? Reversible at low cost points towards auto. Irreversible points towards approve or advise.
  2. Who is affected? Effects confined to the organisation's own housekeeping are different from effects on a customer, an employee or a citizen.
  3. How large is the exposure if it is wrong? Set a limit by value, volume or sensitivity, and route anything above it to approval.
  4. How well is the agent's accuracy known? Without a measured record on this class of action, start at advise or approve.

Answer in that order. A reversible action with a known accuracy record and no effect on individuals is a candidate for auto. Everything else starts higher in the chain and moves down only on evidence.

Where the gate is enforced

A gate that exists only in the user interface is not a gate. The policy engine should enforce it on the decision path, so that scheduled jobs, background tasks and calls from other agents all meet it. Test this by triggering an action class without going through the front end and checking that the gate still holds.

Reading a proposal

Count how many of the proposed use cases fall into each gate. If almost everything is auto, accountability has not been considered. If nothing is auto, the design has not considered value, and the approvers will become a bottleneck that the business works around.

Approval also needs design. An approver who is shown a bare "approve?" button will approve everything. Show the evidence, the alternatives considered and what would change if the action were declined. Track approval times and overrides, because a rising rate of rubber-stamping is as much a control failure as an agent acting without approval.

In practice: List ten actions an agent might take in a process you run. Score each against the four questions and assign a gate. Review the list with the process owner and the risk function, and agree it in writing before development starts. Revisit it quarterly, moving classes down only where the evidence supports the move.

Next: C4, The sealed record, and what reconstructing a decision really requires.

Put it to work

The common mistake. Assigning autonomy to the agent as a whole and enforcing the gate only in the user interface. Scheduled jobs and calls from other agents go around it, and approvers see a bare "approve" button and approve everything.

How to check. Trigger one gated action class without the front end, through a scheduled job or a direct call, and see whether the gate holds. Then review the last batch of approvals and count how many took seconds or changed nothing.

What good looks like. Each action class has a written gate agreed with the process owner and the risk function, enforced on the decision path. Approvers see evidence and alternatives, and approval times and overrides are tracked.

WHICH GATE AN ACTION DESERVESAPPROVEIt matters, but it can be undone.A named person signs before ithappens.ADVISE ONLYIt matters and it cannot beundone. The system prepares; aperson decides.AUTOSmall, and easily reversed. Let itrun, and record that it ran.AUTO, WITH EVIDENCESmall, but it sticks. Run it, andkeep the evidence close.REVERSIBLEHARD TO UNDOHIGHLOWCONSEQUENCETHE HARD LIMITNo agent grants, deniesor alters anentitlement, aneligibility or apenalty. That sitsoutside this grid. It isnot a setting.Autonomy is granted per class of action, never per agent.
Exhibit 02Which gate an action deserves, on consequence and reversibility.

Track C · part 4 of 6

C4. Reconstructing a decision requires the data as it stood, not only a log of what happened

An audit trail answers what the system did. Reconstruction answers why, on what evidence, under whose authority and with what alternatives.

Most systems keep logs. Logs are written for engineers and cover events: a request arrived, a call returned, a job finished. When a regulator, a court or a customer asks why a decision was made, an event log is a starting point and rarely enough.

What reconstruction requires

To rebuild a decision after the fact, you need five things captured at the time.

  • The inputs, as they were. The data the agent saw on that day, or a reference to the exact version of each record. Data changes. If the record has been updated since, a link to the current record shows the wrong picture.
  • The authority. Which principal acted, under which policy version, and who granted that authority.
  • The reasoning path. The retrieved evidence, the tools called, the intermediate results and the model used, including its version and settings.
  • The gate. Whether the action was auto, approved or advised, and if approved, by whom, when, and what they were shown.
  • The outcome. What was done, and what was returned by the downstream system.

Missing any one of these leaves a gap that a sceptical reviewer will find.

What sealed means

A record is sealed when additions are possible and quiet alteration is not. The usual technique is to chain entries. Each entry contains a cryptographic hash of the one before it, so that deleting, inserting or reordering an entry breaks the chain and can be detected. Periodic anchoring of the chain head in a separate system strengthens this.

Sealing matters most with respect to the agent itself. An agent that can edit its own log is unaccountable. The record should be written by the platform on the decision path, with no write access granted to the agent.

Sealing is not the same as encryption or access control. A record can be encrypted and still be edited by an administrator. The test is whether tampering is detectable by someone who does not trust the administrator.

Point-in-time data

The hardest requirement is the second half of the question: can you show the data as it stood when the decision was made? There are two ways to meet it. You can snapshot the inputs into the record, which is simple and increases storage. Or you can use sources with versioned history and store references to the versions used. Either works. Neither is met by pointing to a live table.

Retention and privacy

A complete record may contain personal or confidential data. Plan retention by class of decision, apply the same entitlements to reading the record as to the underlying data, and decide in advance how a deletion request interacts with a record designed to resist deletion. Sealing can protect the structure while individual payloads are redacted or expired under policy, with the redaction itself recorded.

In practice: Choose one decision your organisation made recently with an automated component. Try to reconstruct it from your systems alone, without asking the people involved. List what you could not find. That list is the specification for your record.

Next: C5, Mapping to the frameworks: IMDA agentic guidance, MAS, the EU AI Act.

Put it to work

The common mistake. Pointing to an event log and calling it an audit trail. The log shows that a request arrived and a job ended. It omits the data as it stood, the policy version, the evidence retrieved and what the approver was shown, and its links open live records that have since changed.

How to check. Take one recent decision with an automated component and try to reconstruct it from systems alone, without asking the people involved. List what you cannot find against five items: inputs, authority, reasoning path, gate and outcome.

What good looks like. A decision can be rebuilt as at its date from a record the platform wrote and the agent cannot edit, where tampering is detectable by someone who does not trust the administrator. Retention is set by class of decision.

Requestwhat was asked, by whomEvidencethe records it reliedonPolicy checkthe gate and the ruleApprovalwho, when, what theysawActionwhat was done, whereSealhash of this and thepreviousEach entry carries the hash of the one before it, so any deletion or reordering is detectable.Any decision can be explained from the data as it stood on the day — including which model ran, with which prompt.
Exhibit 03Six parts, hash-chained, reconstructable as of the day.

Track C · part 5 of 6

C5. The three frameworks ask for the same controls, so build them once and map them three times

IMDA, MAS and the EU AI Act differ in legal force and timing, but each expects accountable owners, bounded authority, human oversight and evidence.

This module describes the position as verified in October 2026. Regulatory positions change. Check the primary source before relying on a date.

Singapore: IMDA Model AI Governance Framework for Agentic AI

IMDA announced the framework on 22 January 2026 at the World Economic Forum in Davos. It is guidance, not legislation. It was later updated with case studies and further advice, including on multi-agent systems and third-party agents. It is organised around four dimensions:

  1. Assessing and bounding risks upfront, by choosing suitable use cases and limiting what agents can access and do.
  2. Keeping humans accountable, with defined checkpoints for approval.
  3. Applying technical controls across the agent lifecycle, including testing before deployment and monitoring afterwards.
  4. Enabling end-user responsibility through transparency and training.

Singapore: MAS guidelines on AI risk management

MAS issued a consultation paper on proposed Guidelines on Artificial Intelligence Risk Management on 13 November 2025. The consultation closed on 31 January 2026. The proposals apply to financial institutions and cover generative AI and AI agents. They expect board and senior management oversight, an inventory of AI use, a risk materiality assessment, and controls across the lifecycle, scaled to risk. MAS has also published an AI risk management toolkit through Project MindForge. We have not verified that the guidelines have been finalised, so we give no effective date. Financial institutions should check MAS for the current status.

European Union: the AI Act and the Omnibus

The prohibitions have applied since 2 February 2025, and obligations for general-purpose AI models since 2 August 2025. The Digital Omnibus on AI was published in the Official Journal on 24 July 2026. It defers the high-risk obligations. Stand-alone high-risk systems listed in Annex III now apply from 2 December 2027. High-risk systems embedded in regulated products under Annex I now apply from 2 August 2028. Transparency obligations under Article 50 began on 2 August 2026, with a grace period to 2 December 2026 for marking content from generative systems already on the market.

One control set, three mappings

ControlIMDA dimensionMAS directionEU AI Act theme
Agent inventory and ownerHuman accountabilityAI inventory, oversightProvider and deployer duties
Bounded authorityAssessing and bounding riskRisk-scaled controlsRisk management system
Action gatesHuman checkpointsHuman oversightHuman oversight
Testing and monitoringTechnical controlsLifecycle testingAccuracy and monitoring
Sealed recordTraceabilityDocumentationLogging and record-keeping

In practice: Build your control set from the capabilities in C1 to C4, then maintain a mapping table that cites each framework's own wording. When a date moves, you update the table and not the controls. Ask legal counsel which instruments bind you, since that depends on where you operate and what you do.

Next: C6. The harness is the product: what sits around the model decides whether it can be trusted.

Put it to work

The common mistake. Running a separate compliance project for each framework, each with its own inventory and evidence. The same control is documented three times, and when a date moves the controls are reopened, not just the paperwork.

How to check. Take five controls: inventory, bounded authority, action gates, testing and sealed record. Write beside each the exact wording from every framework that applies to you. Mark any control with no evidence you could show tomorrow, and confirm with counsel which instruments bind you.

What good looks like. One control set and one mapping table that cites each framework's own wording and the date it was verified. When a date or text changes, the table is updated and the controls stay as they are.

WHAT IS ALREADY IN FORCE, AND WHAT IS COMINGFeb 2025EU AI ActProhibited practicesapplyAug 2025EU AI ActGeneral-purpose modelduties applyJan 2026SingaporeIMDA agentic AIframework publishedAug 2026EU AI ActTransparencyobligations applyDec 2027EU AI ActAnnex III high-riskduties, deferredAug 2028EU AI ActAnnex I high-riskobligationsSingapore's framework is guidance and is not binding. MAS guidelines for financial institutions were consulted on and should be checked for their current status.Dates move. This exhibit carries the date it was last checked, and we update it.
Exhibit 04What is already in force, and what is coming.

Track C · part 6 of 6

C6. The harness is the product: what sits around the model decides whether it can be trusted

A production agent is a loop, not a model. The tools, budgets, gates and records around the model are what make it survivable — and they are what you should evaluate when buying.

Every agent demo shows the same thing: a capable model, a goal, and a loop that runs until the goal is met. Production agents are made of something else. Around every model that works for a living sits a harness — the structure that decides what the model can touch, what it can spend, and what happens when it is wrong. The major agent frameworks converged on this shape through 2025 and 2026. The model is rented; the harness is where your organisation lives.

The harness matters because the model is the interchangeable part. Model capabilities converge and prices fall, and a system designed around one model's personality gets rebuilt every eighteen months. A system designed around a harness absorbs a new model the way a factory absorbs a new machine: the guards and procedures stay, only the machine is swapped.

Five layers do the work.

  • Tools with budgets. The agent acts through tools — read a ticket, query a ledger, post a draft. Each validates its inputs, makes its side effects idempotent, and carries a cost. "The agent did it seventeen times" is a budgeting failure before it is anything else.
  • Memory that is inspected. Agents need memory of what was tried and corrected. Memory that accumulates unsupervised becomes folklore, so corrections enter under review and are never promoted into behaviour without a named person's approval.
  • A sandbox and a blast radius. Risky tools run restricted, actions are allowlisted per environment, and anything irreversible waits for a human. Design for the agent's worst day, not its average one.
  • Gates on the decision path. Authorisation lives at the single executor every effect passes through — not in route middleware that a scheduled job can walk around. Each action class is Auto, Approve or Advise, decided before anything is built (module C3).
  • Records and evaluation. Every lap of the loop leaves a signed record (module C4), and an evaluation loop measures the agent's work against outcomes. An agent you cannot observe is an agent you cannot debug; one you cannot reconstruct is one you cannot defend.

The buyer's version of this module is short. When a vendor demos an agent, ask to see the harness instead: where the tool budgets live and what enforces them, the signed record of one action end to end, what happens when it loops at 3am, which actions need your people's approval and where that is enforced, and what breaks if the model is swapped. Good answers to those questions are a product. Good answers about the model are a demo.

In practice: Pick one agentic system you already run or are evaluating. For each of the five layers, write down where it lives and who can change it. Any layer answered with "the framework handles it" or "the model is instructed to" is a layer that does not exist.

Next: D1. Sovereignty is a property of where data, inference and evidence live, not of where the server sits.

Put it to work

The common mistake. Evaluating an agent on its demo — the model, the goal, the loop — and discovering in production that there is no tool budget, no sandbox, and no record an auditor would accept.

How to check. For one agentic system you run or are evaluating, write down where each of the five layers lives and who can change it: tools with budgets, inspected memory, sandbox and blast radius, gates on the decision path, records and evaluation. "The framework handles it" means the layer does not exist.

What good looks like. Each layer exists outside the model, is owned by a named person, and survives a model swap without being rebuilt. The record of one action can be produced end to end, on request, in minutes.

FIG. 1 A PRODUCTION AGENT IS A LOOP, NOT A MODELTHE MODELrented, swappablePOLICY & BUDGETper-agent spend counters and hard limits,enforced outside the modelGATESauto / approve / advise per action class,decided before anything is builtHARNESStools with budgets, memory that isinspected, a sandbox, an evaluation loopEVERY LAP OF THE LOOP LEAVES A SIGNED RECORD
Exhibit 05A production agent is a loop, not a model — the layers that make it survivable.

Something here you disagree with?

These are written to be argued with. If a part of this is wrong, or missing the case you care about, tell us and we will fix it.

hello@eigenforgelabs.ai

Send opens your email client with the note already addressed to us — nothing is stored on this site, and the message goes from your own mailbox, so our reply lands in yours.