EigenForge AI Labs Talk to us
Governed data — track artwork

Learn · Track A

Governed data

Where the answers come from, and why one number matters more than better dashboards.

Track A · part 1 of 5

A1. No single system owns the question, so someone has to assemble the answer

Every system of record holds part of the answer to a business question, and none of them holds all of it.

Take a plain question. Which customers are paying late and also have open quality complaints on goods shipped last quarter? Receivables sit in the finance system. Complaints sit in a service platform. Shipments sit in the ERP or a logistics tool. The customer record sits in a CRM. Each of these systems is correct about its own domain.

None of this is a design fault. Each system was built to run one process well, with its own identifiers, its own calendar and its own idea of what a customer is. The question crosses boundaries that the systems were never asked to cross, and no vendor's roadmap will change that, because the boundary is between your systems and not inside any one of them.

So the work falls to people. Someone exports from each system, reconciles the results in a spreadsheet and assembles the answer. The pattern has three predictable consequences.

  • Timing. The assembly happens periodically, so the answer arrives after the decision needed it.
  • Copies. Each function builds its own version, with slightly different filters, dates and exclusions. The review meeting then debates the data and not the decision.
  • Dependence. The knowledge of how to assemble the answer lives in one or two people, and it leaves when they do.

There are three broad ways to bring the parts together. You can copy everything into a central store, such as a warehouse or lake. You can query the systems where they are and join the results at question time. Or you can combine the two, copying what must be copied and querying the rest. Later modules in this track cover the trade-offs. For now, the point is that the choice of technology comes second. The first task is to describe the question precisely.

A usable description has four parts. Name the question in one sentence. List the systems that hold each part of the answer. Identify the key that joins each pair of systems. State who is allowed to see the result. If you cannot name the join key between two systems, you have found the real project, and it is a matter of identity and not of tooling. Customer numbers that differ between billing and service, or part numbers that were reissued after an acquisition, will defeat any platform you buy.

You can recognise the problem without any technical review. Two people ask the same question and receive different answers. A month-end close needs several days of reconciliation. The count of active customers depends on which team you ask. Any of these signals a question that no system owns.

It helps to say what the problem is not. It is not a shortage of data, and it is rarely a shortage of analysts. Most organisations hold more than enough data. They lack an agreed route from the question to the records that answer it.

In practice: Pick one recurring question that currently takes a person more than an hour to assemble. Write down the systems, the join keys and the people who may see the result. Then ask who owns the answer today. If the honest reply is that nobody does, you have identified your first candidate for a governed layer.

Next: A2, Reading systems in place avoids migration, and it costs load, latency and discipline.

Put it to work

The common mistake. Teams choose a platform first and look for questions to put on it afterwards. The join between systems, such as a customer number that differs between billing and service, turns up halfway through the build, when it is costly to fix and the schedule is already committed.

How to check. Pick one recurring question. Ask the person who assembles it to show every export and every manual matching step. Count the hand-edits made to keys such as customer or part numbers. Each one is an undocumented join, and together they are your real project.

What good looks like. One named owner for the question, a written list of systems and join keys, and a reproducible answer. Two people ask and receive the same figure without a private spreadsheet in between.

ERPcorrect aloneShop floorcorrect aloneLab & testcorrect aloneCRMcorrect aloneMaintenancecorrect aloneHR & payrollcorrect aloneSpreadsheetscorrect aloneNo system owns the question.Every one of them owns part of the answer.The work goes to peopleExport, reconcile, assemble — periodically, becausedoing it continuously would consume the team.Every function keeps its own copyCurated for its own purpose, defensible on its ownterms, quietly different from everyone else's.
Exhibit 01Seven systems, each correct on its own. The work of joining them falls to people.

Track A · part 2 of 5

A2. Reading systems in place avoids migration, and it costs load, latency and discipline

Querying source systems directly removes the second copy of the data, and it moves some of the cost to the sources.

Reading at source means that a query reaches the system of record, or a replica of it, at the moment someone asks. The data is not first extracted, transformed and loaded into a new store. Vendors describe this as federation, virtualisation or query-in-place. The common feature is that the source remains the only copy.

The advantages are real. There is no second copy to drift away from the source. There is no migration programme to fund and sequence. If the access is read-only, nothing can write back, nothing changes in transaction processing, and a security reviewer has a much shorter list of risks to consider. A useful test for any such design is to switch it off and see whether everything carries on exactly as before.

The costs are also real, and a buyer should ask about each of them.

ConcernWhat to ask
Load on the sourceDo analytical queries compete with transactions? Can reads go to a replica, be scheduled or be limited?
LatencyHow long does a join across two remote systems take, and where is the join executed?
AvailabilityIf the source is down or in a maintenance window, what does the user see?
HistoryDoes the source keep the history you need, or does it overwrite?
Schema changeWhat happens to queries and definitions when a source upgrade renames a field?

The last two deserve emphasis. Many operational systems overwrite old values, and reading them tells you only how things stand now. If a question needs history, you must either capture changes as they occur or accept that the history does not exist. Schema change is the quiet cost. A source upgrade can break a join without any error, and the break appears as a wrong number.

Reading at source also does nothing to improve the quality of the records. Duplicate customers and inconsistent codes remain in the source. The quality rules and entity matching must therefore run at read time, or the problems appear in every answer.

There are good reasons to copy. Heavy, repeated aggregation over very large tables belongs in a store built for it. Some sources cannot tolerate any extra load. Some history must be preserved because the source will not keep it. A sensible architecture is mixed, and the useful discipline is to copy for a stated reason and not by default. EigenForge's ZigmaData platform, which is in production, takes this approach: it reads systems of record where they sit, read-only, and it treats an existing warehouse as one more source and not as something to replace.

In practice: For your priority question, mark each source as read directly, read from a replica or copy, using the table above. For every source you mark as a copy, write down the reason. A copy with no stated reason is a future reconciliation problem.

Next: A3, One agreed number matters more than another dashboard.

Put it to work

The common mistake. Adopting "no second copy" as a principle and then pointing analytical queries at the live transactional database. The first month-end run slows order entry, operations withdraws access, and the project is blamed for a load decision nobody wrote down.

How to check. Ask the owner of each source system two things: when is it busiest, and how much read load can it tolerate? Then run your heaviest expected query against a replica or in a quiet window, time it, and watch the source's own monitoring while it runs.

What good looks like. Each source is labelled as read directly, replica or copy, with a written reason for every copy. Heavy queries run off peak or on replicas, and operations can see and limit the load.

Track A · part 3 of 5

A3. One agreed number matters more than another dashboard

A dashboard displays a figure. Only an agreed definition, a traceable lineage and a quality rule can settle an argument about it.

Ask three teams for last month's revenue. Finance may report recognised revenue. Sales may report orders booked. Operations may report goods shipped. All three are defensible and all three are labelled revenue. The meeting that follows spends its time on which number is right, and the decision waits.

Building another dashboard does not help, because the dashboard inherits whichever definition its author chose. It adds a fourth presentation of the same disagreement. The remedy lies underneath the dashboard, in three things that are easy to name and hard to maintain.

ElementThe question it answersWhat goes wrong without it
DefinitionWhat exactly does this measure count?Teams report different figures under one name
LineageWhich source records and versions produced this figure?Nobody can check it, so nobody trusts it
Quality ruleIs the input fit to be counted?Bad records enter silently and distort the total

A definition is a business decision, and it needs a named owner. It states what is included, what is excluded, the date basis and the currency treatment. It is published where people can read it, and it is applied in the layer that computes the measure. If each report re-implements the definition, the definitions will diverge within months.

Definitions also change. When finance revises how a measure treats returns, the history before and after the change is not comparable unless the definition is versioned and the change is recorded. A figure without its definition version is a figure that cannot be compared with last year.

Lineage means that a figure resolves to the source system, the object, the row and the version it came from. This is what turns an analytical answer into one an auditor will accept. Without it, a challenge to a number is settled by whoever argues most confidently.

Quality rules work best when written in business language and applied as the data arrives, with failures reported as exceptions. Silent correction is a hazard. If a system quietly fixes a bad record, the fix is invisible, the cause is never addressed and the number cannot be reproduced from the source. A visible exception can be assigned to someone and closed.

There is a cultural point as well. Agreeing a definition forces a conversation that many organisations have avoided for years. The conversation is uncomfortable, and it is the actual work. The technology only records and enforces the outcome.

In practice: Choose the three measures that appear most often in your leadership meetings. For each, find out how many distinct definitions are in use, and who could decide between them. Publish the chosen definition with a version and an owner before anyone builds a further report.

Next: A4, Entitlement belongs at query time, because the front door is not the only way in.

Put it to work

The common mistake. Publishing a definition on a slide or wiki page while every report keeps computing the measure with its own formula. Within a few months the written definition and the reports disagree, and nobody knows which one drifted.

How to check. Take one headline measure and ask three report authors to show the formula they use. Compare exclusions, date basis and currency treatment line by line. Count the differences, then ask who has the authority to settle them.

What good looks like. The measure has one owner and one versioned definition, applied in one place, and every report reads from it. A changed definition appears as a dated version, so year-on-year comparisons say which rule they use.

TODAYWITH A GOVERNED LAYERA cross-system questionDays of manual assemblyAsked and answered, with lineageThe management packRebuilt every periodProduced from governed definitionsA historical recordRetrieved from archive, if at allQueried in seconds, with its sourceA quality or claim disputeSettled for want of evidenceAnswered with provenanceA containment decisionScoped wide to be safeScoped to what is actually affectedA routine caseRead, routed and drafted by handPrepared by an agent, decided by a personAn AI initiativeStalled on data readinessGrounded on governed recordsA new reporting requestA ticket and a queueSelf-service in plain languageNone of the right-hand column requires a system to be replaced.
Exhibit 02What changes in practice when definitions are agreed once.

Track A · part 4 of 5

A4. Entitlement belongs at query time, because the front door is not the only way in

Checking who may see data when they log in protects the application. Checking when the query runs protects the data.

The front-door model is familiar. A user signs in to an application or a BI tool, and the tool decides which screens and reports that person may open. This works while every route to the data passes through that one tool. In a real estate it never does.

Consider the other routes. A scheduled job extracts a table to a shared folder. An analyst exports a report to a spreadsheet and sends it on. A dashboard connects with a shared service account that can read everything, and filters the display afterwards. A chat assistant is pointed at the same tables. A background process reads data on a timer and never signs in at all. Each of these bypasses the front door, and the control that sat there does not follow the data.

Query-time entitlement moves the check to the point where the data is read. When a query executes, the layer evaluates the identity of the person asking, together with attributes such as role, region or clearance, and applies policy to that query. Row-level policy removes the rows the person may not see. Column-level policy masks or removes fields such as salary or national identifier. Two people who ask the same question receive answers limited to what each is entitled to, from the same layer, and no separate extracts exist to be kept in step.

The approach has costs, and you should plan for them.

  • Policy modelling. Someone must translate the rules in policy documents into conditions the system can evaluate. This is slow the first time, and the ambiguities in the written rules surface here.
  • Identity. Policy depends on a trustworthy identity for every caller, including services and agents. Using the existing directory avoids a second identity store to administer.
  • Performance. Evaluating policy on every query adds work. Ask how the overhead behaves on your largest tables.
  • Consistency across paths. The same policy must apply to SQL, to BI tools, to natural-language questions and to agents. A path that bypasses the check makes the rest decorative.

Testing is simple to describe. Pick two users with different entitlements. Ask each the same set of questions through every access path. Compare the results to what policy says each should see. Repeat after every policy change, because a change that fixes one rule often breaks another.

Audit completes the picture. A record of who queried what, and when, lets you answer an access question after the event. Without one, you can assert that policy was enforced, but you cannot show it.

In practice: List every route by which data in your priority system currently reaches a person or a process, including exports, service accounts and scheduled jobs. For each route, ask whether the person's own entitlement is checked when the data is read. Any route where the answer is no is a gap in the control, however good the front door is.

Next: A5, Explaining a past decision requires the data, the definitions and the policy as they were that day.

Put it to work

The common mistake. Securing the BI tool and assuming the data is secured. The dashboard connects through a shared service account that can read every row and filters the display afterwards, so any export, notebook or chat assistant on the same account sees everything.

How to check. List the service accounts and scheduled jobs that read your priority system. For each, ask whose entitlement applies when it runs. Then have two users with different rights ask the same ten questions through every route and compare the results.

What good looks like. The same question asked by two people returns two answers, each limited to that person's rights, through SQL, BI tools and natural language alike. Every policy change is followed by a rerun of the two-user test.

Track A · part 5 of 5

A5. Explaining a past decision requires the data, the definitions and the policy as they were that day

A system that can only tell you how things stand now cannot account for what it told someone last March.

An auditor, a regulator or a customer in dispute asks a specific question. What did the system show, and on what basis, on a named date? Answering requires more than a backup. It requires that several different things were kept, and that they can be put back together.

Start with the two kinds of time. Valid time is when a fact was true in the world. Transaction time is when the system learned about it. Suppose an invoice for a February sale was corrected in April. The February report, as it was produced in March, used the original invoice. A restated report uses the corrected one. Both are legitimate, and a system must be able to produce either on request. Tracking both clocks is often called bitemporal storage, and most operational systems do not do it, because they overwrite.

Five things must be retained to reconstruct a decision.

  • The data. Prior values, not only current ones, with the time each was recorded.
  • The definitions. The version of each measure that applied on the date, since definitions change.
  • The policy. The entitlement rules in force, so you can show what the person was allowed to see.
  • The evidence. For an AI answer, the records and passages retrieved, not merely the final text.
  • The model and its settings. Which model ran, with which prompt and configuration, because a different version may answer differently.

Retention alone is insufficient if the record can be changed afterwards. A tamper-evident log, in which each entry carries a hash of the one before it, makes deletion or reordering detectable. The test is not whether logs exist. It is whether you can hand an auditor an account of what happened and have it hold up.

Cost needs attention. Keeping every version of every record is expensive, so decide which data and which decisions carry this obligation. Statutory retention rules often give a starting point. Tiered storage, where older data sits on cheaper media but remains directly queryable, avoids the restore step that turns an audit request into a project.

Few vendors can demonstrate this end to end. EigenForge's Kyros platform, powered by Kaman and in production, produces a sealed record per agent action and can reconstruct a decision as of its date, and the same questions are fair to put to any alternative.

In practice: Choose one decision your organisation made six months ago on the basis of a report or an AI answer. Try to reproduce exactly what was shown then, with the same data, definitions and access rights. Where you fail, note which of the five items was missing. That list is your retention requirement.

Next: B1, Pilots demonstrate well because the conditions that cause failure are absent.

Put it to work

The common mistake. Treating backup as audit. Teams keep current data and event logs, then discover that the definition version, the access rules and the retrieved passages from the date in question were never stored, so the answer cannot be rebuilt.

How to check. Choose a decision made six months ago and try to reproduce the report or AI answer exactly as it was shown. Tick off data, definitions, policy, evidence and model settings. Every missing tick is a retention requirement, so write it down.

What good looks like. For decisions that carry the obligation, you can produce the data, definition version, policy, evidence and model settings as at the date, held in a tamper-evident log that an auditor accepts.

IdentityYour existing directory. No separatestore.AuthorisationRow and column policy at queryexecution.EncryptionIn transit and at rest, keys under yourcontrol.AuditEvery query, access and agent action.LineageTo source record and version.ResidencyPer jurisdiction, with a federated view.RedactionBy jurisdiction and by entitlement.AI boundaryInference inside your environment.Security and governance enforced at the point of use, in every topology.
Exhibit 03Where each control is enforced, and at what moment.

Something here you disagree with?

These are written to be argued with. If a part of this is wrong, or missing the case you care about, tell us and we will fix it.

hello@eigenforgelabs.ai

Send opens your email client with the note already addressed to us — nothing is stored on this site, and the message goes from your own mailbox, so our reply lands in yours.