
Research
We did not start with a product. We started with ten problems.
Everything we build began as something a customer could not do. We work on the problem first, usually for longer than is commercially sensible, and the product is what is left standing afterwards. Eighteen areas of work, each carrying what has been built and where you can get it.
The ten problems
Six came from clients. Four came from work we could not do.
No system owns the question
Ask the question the business has, not the one a single system can answer.
Confident answers nobody can check
It shows its working, or it says it cannot.
Agents nobody can hold responsible
Agents work like employees: a role, an owner, a budget, a record.
Your context living in someone else's system
Own your context. Rent your intelligence.
Programmes that cost more than they return
Value early, investment following evidence, a declining cost per question.
No usable answer to what it actually did
A record you hand to an auditor, not a log you have to explain.
Data that is not allowed to leave
Sovereignty as an ordinary configuration, not a diminished edition.
Systems nobody can get data out of
If a program can produce it, it can be queried.
Every function reporting a different number
One set of definitions, agreed once, enforced everywhere.
Proving the target equals the source
Reconciliation on live data across both estates, with nothing written back.
Why this agenda, and why now
The market has caught up with the problems.
Through 2025 and 2026 the industry's own surveys converged on an uncomfortable picture: most organisations are piloting agents, only a small minority have them in production, and Gartner has projected that more than two in five agentic AI projects will be cancelled by the end of 2027 — citing escalating cost, unclear value and inadequate risk controls. Not model quality. Data quality is named as the top blocker more often than anything else.
Read the eighteen areas below against that list and the correspondence is not a coincidence. The programmes that stall are not failing at intelligence; they are failing at grounding, at evidence, at cost control and at accountability — the unglamorous disciplines this page is about. Meanwhile the ground is shifting under two other assumptions: inference is moving decisively toward private infrastructure as context becomes the asset worth protecting, and token economics has become a board-level line item rather than an engineering detail. Both are areas of work here, and both are written up in plain language in the Learn tracks and the blog.
How to read this page
Two different things, tracked separately.
Whether the work is built, and where you can get it. These are not the same. A capability can be implemented and tested and still sit inside a product that is not generally available, which is the position most of the list below is in.
Of the eighteen areas: four run in client deployments today, six are built and with design partners, five are built and running on our own estate with no external users, two are in development, and one is proposed as the next build.
Group
Reading and answering
How a question reaches an answer at all. All four run in client deployments today.
Area 01 · against problems 01, 02
Grounding an answer in records, not in a copy of them
In production In production, on the data platform we deliver on.
Answers come from your live records, so the number in the report is the number in the system — not a copy of it from last Tuesday.
The technical detail
The common approach embeds documents into a vector store and retrieves by similarity. We took the view that this loses what enterprises most need, which is the relationships between records, and that answers drift from their source as the copy ages. The work went into resolving a natural-language question against an ontology built over live records and generating a query, instead of retrieving passages that look relevant.
Area 02 · against problems 01, 06
Entitlement decided at the moment of the question
In production In production.
Two people ask the same question and each correctly sees only what they are allowed to see — enforced when the answer is computed, not when they logged in.
The technical detail
Most systems check permission at the front door and then trust the session. That fails as soon as one question spans several systems with different permission models. The work was on evaluating row- and column-level entitlement per user at query execution, so two people asking the same question correctly receive different answers.
Area 03 · against problems 08
Reading a system that has no interface
In production In production, and the subject of a patent filing.
The system nobody can get data out of becomes readable without being replaced — the green screen answers back.
The technical detail
Some systems have no API, no supported export and nobody left who wrote them. The conventional answer is a rewrite. The work produced a method that treats a program's output as a queryable table, so the system can be read without being changed or replaced.
Area 04 · against problems 01, 04, 07
Making a question answerable without moving the data
In production In production.
Nothing moves before value is visible: your systems keep running, we read them where they sit.
The technical detail
Reading several systems in place, joining them, and keeping the source system as the authority throughout. The research was mostly about what not to do: where landing data for performance is unavoidable, and how to keep lineage intact when it is, so no derived figure ever becomes the record.
Group
Evidence, authority and record
Whether the answer can be trusted, and whether anyone can be held to it. Built, and with design partners or on our own estate.
Area 05 · against problems 02
Refusing outright instead of hedging
Private beta Built and enforced. In private beta with design partners.
When the evidence is not there, the system says so — a recorded refusal instead of a fluent guess.
The technical detail
An AI system that produces a fluent answer when the material does not support one is more dangerous than one that fails loudly. Citation attachment and confidence hedging are widespread; refusal as the enforced default outcome is not. An answer that fails its evidence checks is not returned at all, and the refusal is persisted with a machine-readable cause, which means the refusals accumulate into a dataset in their own right.
Area 06 · against problems 02
Claims tied to the passage that supports them
Private beta Built. In private beta.
Every sentence in an answer is one click from the record it came from, so checking takes seconds, not an afternoon.
The technical detail
A citation that exists is not the same as a source that supports the claim. The work attaches each answer to the specific passages and records it was generated from, so a reader can go from a statement to the underlying row in one step, and so an unsupported statement is caught by the evidence gate above instead of being softened.
Area 07 · against problems 03, 06
Enforcement at the executor, not at the transport
Private beta Built. In private beta.
The same rules govern an agent whether it arrives by web page, queue or scheduled job — there is no back door for automation.
The technical detail
The standard pattern puts authorisation in route middleware, which governs the web request and nothing else. Anything reaching the system by another path, a scheduled job, a queue consumer, an internal call, is governed by a different rule or by none. The work moved enforcement to the single executor through which every effect is performed, so non-web execution paths are governed identically and an unregistered action fails closed.
Area 08 · against problems 03
Autonomy granted by action, not by agent
Private beta Built. In private beta. We are mapping it against the agentic AI governance framework Singapore's IMDA published in 2026 and will publish that mapping.
You grant autonomy per action, not per agent: routine things flow, irreversible things wait for a named person.
The technical detail
Giving an agent a permission level is the wrong shape. Autonomy is granted per class of action, judged on consequence and reversibility, and enforced on the decision path. Underneath it sits a hard limit: no agent grants, denies or alters an entitlement, an eligibility or a penalty.
Area 09 · against problems 03, 06
A record an auditor can use
Private beta Built. In private beta.
An auditor gets a reconstruction of what was decided, on what basis, under whose authority — not a log file to interpret.
The technical detail
A log tells you what happened. An auditor needs to know what was decided, on what basis, under whose authority, and what the data looked like at the time. The records are hash-chained and the ledgers are append-only, enforced by the storage layer itself and not by application code, so a deletion or an edit is detectable rather than trusted not to happen.
Area 10 · against problems 03
Agents as principals in the authority model
Internal testing Built. Running on our own estate. Not yet externally exercised.
Agents appear in the same authority and audit model as people: an owner, a budget, a record of their own.
The technical detail
Agent frameworks are numerous. Agents registered as principals in the same authority, budget and audit model as people is rare. An agent has an identity, an owner, a scope and a spend counter that cannot be raced by concurrent requests, and it appears in the same record as a human actor doing the same thing.
Area 11 · against problems 03, 06
A governance property enforced at build time
Internal testing Built and enforced in our own pipeline. Not yet offered as a product.
A governance guarantee that depends on someone remembering a test is not a guarantee — our build refuses code that cannot prove its record fired.
The technical detail
Coverage gates are common. A gate that fails the build when a governed capability arrives without an executable check proving its record actually fired is not, as far as we can find. The point is the inversion: a governance guarantee that depends on somebody remembering to write a test is not a guarantee, so the build refuses the code instead.
Area 12 · against problems 07
Sovereign deployment as configuration, from one codebase
Private beta Built. Validated in development and with design partners, not yet in a production deployment.
An air-gapped site runs the same product as a connected one, not last year's edition with an apology.
The technical detail
On-premises variants are common. Structural parity, where the disconnected deployment is the same code as the connected one and differs only by configuration, is not. That is what stops an air-gapped client being sold last year's product.
Group
Measurement and cost
Knowing whether it is working, and what it costs to keep working.
Area 13 · against problems 02, 05
Retrieval lanes as data, against a pinned corpus
Internal testing Built. Running on our own estate.
We can change one part of the retrieval pipeline at a time and prove what it did — improvement by measurement, not by release notes.
The technical detail
Configurable retrieval pipelines exist everywhere. Coupling them to an evaluation harness with a versioned corpus, so one variable can be changed at a time and the comparison means something, is the specific contribution. A retrieval lane is authored data, not code, which is what makes the comparison cheap enough to actually run.
Area 14 · against problems 05
Routing to the cheapest model that passes
Internal testing Built. Running on our own estate and used as a method in delivery today.
Each question goes to the cheapest model that passes your own evaluation — the bill falls without the quality moving.
The technical detail
Most AI cost conversations are about price per token, which nobody can act on. Each workload is routed to the smallest model that passes the customer's own evaluation, measured as cost per answered question including the cost of being wrong, from a per-call ledger that carries cost, latency, tokens and a difficulty score for every request.
Area 15 · against problems 02, 05
Shadow evaluation alongside production
Internal testing Built. Running on our own estate.
A change is measured against real traffic before it is adopted, so a regression is a divergence we catch, not a complaint you file.
The technical detail
An evaluation path runs beside the live one without affecting what the user receives, so a change can be measured against real traffic before it is adopted, and a regression shows up as a divergence instead of as a complaint.
Area 16 · against problems 02, 05
Learning from corrections without learning the wrong thing
In development In development.
When your people correct an answer, the system gets better in a way that stays inspectable — and reversible.
The technical detail
When a person corrects an answer, that correction is the most valuable signal the system will ever receive and the most dangerous. Corrections go into the knowledge layer under review, never into the model, so an improvement stays inspectable and reversible. The refusal set and the decision log, which records rejected experiments alongside adopted ones, are the raw material.
Area 17 · against problems 01, 05
Separating what the engine cannot do from what the corpus does not contain
In development In development. An annotation protocol exists; the automated separation does not.
When an answer is poor we can tell whether the engine is weak or the material is — the two need opposite fixes.
The technical detail
When an answer is poor, retrieval and generation may have been weak, or the answer may simply not have been in the material. These need opposite fixes and look identical from outside. Every enterprise AI post-mortem we have seen gets stuck here, and teams rebuild the engine when the corpus was the problem.
Group
Looking ahead
One item — the next thing we intend to build, and why we think it matters beyond us.
Area 18 · against problems 02, 05
Invariant checking for a system degrading silently
Proposed Proposed, with motivating incidents documented.
The dangerous failure is the silent one — answers that still look fine after something upstream changed. This is the check we build next.
The technical detail
Loud failures are easy. The dangerous case is a schema change upstream, a model swap or a stale index, where the answers still look fine. The idea is a runtime invariant that fails when a governance property stops holding, in the way a type system fails when a contract is broken.
Over to you
What should we be working on?
Several of the eighteen areas above exist because somebody told us about a problem we had not thought about. If there is something in your organisation that nobody has solved, not a requirement for a product but a problem you believe is genuinely unsolved, we would like to hear it.
We read all of them. We will tell you plainly whether we think it is tractable, whether someone has already solved it, or whether we have no idea. There is no obligation in either direction and no sequence behind this form.
The question your organisation cannot answer, and that you suspect nobody can
The thing everyone has accepted is impossible, and quietly works around
The control your auditors ask for that no system can produce
Tell us what to work on
A few sentences is plenty. What you have already tried is the most useful part.
What informs the work
What we read.
Each mapped to the area of work it bears on.
| Area | The work | Why it matters to us | How we use it |
|---|---|---|---|
| Grounding annotation and evaluation | Benchmarks that annotate which span supports which claim | Bears directly on area 11 — grounding labels are what a serious benchmark needs. | On the list to run against. |
| Source attribution in agents | Work parsing and evaluating whether cited sources actually support the claim | The exact failure the evidence gate is designed against: a citation that exists but does not support the text. | Implemented in the evidence gate. |
| Graph, vector and hybrid retrieval | Comparative analyses of retrieval architectures | The literature our ontology position is tested against. | Built into the platform; reviewed on every engagement. |
| When graph retrieval does not help | Practitioner analyses arguing graph retrieval is over-applied | A counterweight to our own design — we read what argues against us. | The subject of one of our own articles. |
| Agentic AI governance | IMDA's Model AI Governance Framework for Agentic AI, Singapore, 2026 | Our home regulator, on our exact subject. Our action-gate model maps to it. | Mapped in delivery work. |
| Financial-sector AI risk | MAS guidance on AI risk management for financial institutions | Directly relevant to banking and insurance work. | Applied in sector engagements. |
| EU AI Act obligations | The Act as amended, including the deferral of high-risk obligations | Changes what clients must do and when. Affects the sovereign and evidence argument. | Tracked as obligations move. |
| Open-weights serving economics | Cost-per-resolved-query analyses of local inference | Underpins the cost-per-answered-question position. | The method in daily delivery use. |
Working with us
Four ways, none of which needs money to start.
Argue with us
Write and tell us where we are wrong. No agreement and no meeting required. This is the one we most want.
Bring an evaluation
If you have a benchmark, a corpus or a test harness relevant to any area above, we will run our system against it and publish the results.
Place a researcher
A student, a postdoc or an industrial-placement researcher working on one question with access to a real system and real constraints. We would host. We are not offering to fund.
Joint proposal
If a programme exists that fits one of these questions and you need an industry partner, we are willing to be that partner. We would not lead.
Or write directly:
research@eigenforgelabs.ai
To tell us what we should be working on, to send us something we should have read, or to tell us which part of this you do not believe.
Talk to us