EigenForge AI Labs Talk to us
Enterprise AI that holds up — track artwork

Learn · Track B

Enterprise AI that holds up

Why pilots stall, what grounding really requires, and what the bill is actually made of.

Track B · part 1 of 7

B1. Pilots demonstrate well because the conditions that cause failure are absent

A pilot shows that a model can answer. Production requires that it answers correctly, for the right people, at a cost someone has agreed to pay.

Most stalled AI programmes did not fail on model quality. They stalled because the pilot was run under conditions that production will not offer. The demonstration was genuine, and it answered a different question from the one the business later needed answered.

Compare the two settings.

Pilot conditionProduction condition
A curated extract of clean dataLive systems with duplicates, gaps and late updates
A dozen prepared questionsThousands of unprepared ones, including ambiguous and out-of-scope
A few friendly usersMany users with different entitlements
One administrator with full accessPer-user policy on every query
Cost ignoredA budget with an owner
Accuracy judged by impressionAccuracy measured against agreed answers
A sponsor who likes the resultAn accountable owner for the outcome

Each row is a place where the pilot hid a problem. The curated extract concealed the join errors between systems. The prepared questions concealed how the system behaves when it does not know. The single administrator concealed the entitlement problem entirely, and entitlement is often the issue that stops a security review and ends the programme.

There is a sequence behind this. Making data answerable comes first. Making AI usable on that data comes second. Letting agents do the routine work comes third. Programmes that attempt the third stage before the first is complete tend to stall, because an agent acting on data nobody trusts only produces wrong actions faster.

Production also introduces work that a pilot never meets. Someone must monitor answer quality over time as data and models change. Someone must handle the questions that fail. Someone must pay for the compute. These are operating tasks and they continue indefinitely, so a pilot that has no plan for them has only postponed the cost.

You can design a pilot to learn about production instead of to impress. Use real data from live systems, even if the scope is narrow. Include users with restricted access. Write the success criterion down before the work starts, in terms of the business question and not of the technology. Agree in advance that a verdict that the work should not proceed is an acceptable result. A pilot that cannot return that verdict is a demonstration.

Finally, ask who will own the result. A pilot with a sponsor but no owner usually ends when the sponsor's attention moves on.

In practice: Before approving the next pilot, go through the table above and mark each row as it will apply to the pilot. Every row that reads as a pilot condition is a risk you are choosing to carry into the production decision. Fix the cheap ones now, starting with live data and restricted users.

Next: B2, A citation shows where an answer came from, and evidence shows that it is supported.

Put it to work

The common mistake. Writing the success criterion after seeing the pilot results, and phrasing it as "users liked it". The pilot then cannot return a negative verdict, and entitlement, cost and ownership go untested until the production review.

How to check. Read the pilot charter. Is there a criterion in business terms, dated before work started? Does it name an owner other than the sponsor, include a restricted-access user, and state which result means stop? Add whatever is missing before the next sprint.

What good looks like. The pilot ran on live data with restricted users, was judged against criteria written in advance, and ended with a recorded verdict, including the option to stop, signed by a named owner.

1ReadSystems of record, in place. Read-only, nomigration, no write-back.Switch it off and everything carries on.2GovernDefinitions agreed once. Quality onarrival. Lineage on every figure. Policyper user.One answer, per entitlement.3ActAgents that prepare and act inside limitsa named person set.Every action reconstructable.THE TESTSwitch it off — does everything carry on exactly as before?
Exhibit 01The order that matters. Most stalled programmes attempted the third step first.

Track B · part 2 of 7

B2. A citation shows where an answer came from, and evidence shows that it is supported

A link beside a sentence proves only that a document exists, and the reader still has to check whether it says what the sentence claims.

Retrieval-grounded systems are often sold on the strength of their citations. A response appears with a reference to a source, and the reference creates confidence. The confidence is sometimes misplaced, because a citation and a piece of evidence are different things.

A citation names a source. Evidence is the specific passage or record that supports the specific claim, in the version that existed when the answer was given, in a form the reader can inspect. The gap between the two is where errors hide.

There are several common ways for a citation to mislead.

  • The source exists but does not say it. The model produces a plausible sentence and attaches a real document that is on the same topic.
  • The source is right but the figure is invented. The passage supports the general point, and the number in the answer comes from nowhere.
  • The source is stale. The document was superseded, and the system retrieved the old version.
  • The source is real but unauthorised. The user was not entitled to the document, and the answer revealed its content.

For a document corpus, evidence means a quoted passage with its location and version. For structured data, the evidence is different. It is the query that was run and the rows it returned. A fluent sentence about total sales is only checkable if the reader can see the query, the filters applied and the records counted. Showing the generated query alongside the answer is a stronger form of evidence than any link.

Four tests help to separate the two.

  1. Can the reader open the cited passage or rows in one step, without searching?
  2. Does the passage actually entail the claim, which means a reasonable reader would agree the claim follows from it?
  3. Is the version shown the version that existed when the answer was given?
  4. When nothing relevant is retrieved, does the system decline to answer, or does it write something anyway?

The second test can be partly automated. A second check, using a model or a rule, compares each claim to the passage cited for it and flags those it does not support. This check is imperfect, and it should be measured on your own material, as module B5 explains. It still catches a good share of the gross failures, and it gives the reader a signal that a link alone cannot provide.

The fourth test leads into refusal, which is the subject of B4. A grounded system is defined as much by what it will not say as by what it can quote.

In practice: Take ten answers from any retrieval system you are evaluating, including a competitor's. For each, open the citation and mark whether the cited text supports the exact claim. Count how often it does not. That ratio tells you more about the system than a demonstration of its best answers.

Next: B3, Ontology and vector approaches suit different questions, and neither suits all of them.

Put it to work

The common mistake. Treating the presence of a citation as proof. Reviewers glance at the link, see a document on the right topic and approve, so answers with invented figures or superseded sources pass because a real document sits beside them.

How to check. Take ten answers from a retrieval system you use and open each citation. Mark whether the passage supports the exact claim, whether it is the current version, and whether the user was entitled to read it. The share that fails is your first measure.

What good looks like. Each answer shows the quoted passage, or the query and rows, in the version that existed at the time, and the reader can open it in one step. When nothing relevant is found, the system says so.

THE COMMON APPROACHChunk, embed, retrieve bysimilarity✕Parent–child relationships are lost✕Joins cannot be rebuilt from similarity✕The copy drifts from the source✕Fluent, confident, sometimes wrong✕Content leaves the boundary to be embeddedAN ONTOLOGY OVER LIVE DATAModel relationships, resolve,query✓Table relationships preserved, joins correct✓Answers generated from live records, never from a stale copy✓Every answer traces back to records✓Policy applied at query time, per user✓Runs against a local model — nothing leavesIs it answering from a copy of your data, or from your data?
Exhibit 02The distinction that decides whether an answer can be trusted.

Track B · part 3 of 7

B3. Ontology and vector approaches suit different questions, and neither suits all of them

Vector retrieval is the better choice for text and for speed to a first version, and an ontology is the better choice for records that must be joined and counted.

Two retrieval designs dominate enterprise AI. They solve different problems, and arguments that treat one as superior in general usually come from selling it.

Vector retrieval splits content into chunks, converts each chunk to a numerical embedding and, at question time, returns the chunks nearest to the question in meaning. An ontology approach models entities and their relationships explicitly, such as customers, orders and shipments, resolves a question against that model and generates a query over the live records.

Vector retrieval is the better choice in these circumstances.

  • The corpus is mostly documents: policies, contracts, manuals, emails, reports.
  • The question is about what a text says, or about finding material that resembles another.
  • The schema is unknown, unstable or absent, so there is nothing to model.
  • You need a working first version quickly, with little preparatory effort.
  • The vocabulary is messy, and meaning-based matching tolerates synonyms and spelling that an exact model would not.
  • The stakes are low enough that a reader will check the output.

An ontology is the better choice when the answer depends on structure. Relating a customer to their orders and then to their shipments is a join, and a similarity search over chunked tables cannot reliably reconstruct it. Counting, summing and filtering with exact conditions are also queries, not matters of resemblance. A generated query is either correct or fails visibly, and every figure traces to records. Policy can be applied when the query runs.

An ontology has its own weaknesses, and they are significant.

  • Building and maintaining the model takes effort and needs an owner. If nobody maintains it, it decays.
  • Concepts that were not modelled cannot be asked about. The system fails, which is visible but still a failure.
  • It does little for free text. A contract clause is not a row.
  • The first useful answer takes longer than it does with a vector index.
Question typeUsually the better fit
What does our policy say about travel approval?Vector
Which clauses in these contracts resemble this one?Vector
How many orders shipped late to customers with open complaints?Ontology
What was gross margin by site last quarter?Ontology
Summarise the issues raised in these service notes, then count the affected accountsBoth

The last row is the common case. Real questions often mix a count over records with an explanation drawn from text, and hybrid designs use each approach for what it does well. When you evaluate a product, ask which parts of your question set are records and which are text, and test each part separately.

In practice: Sort your fifty most likely questions into three piles: about text, about records, and both. The proportions between the piles tell you where the weight of the system should lie, and which vendor claims apply to you.

Next: B4, A system that never refuses tells you nothing when it answers.

Put it to work

The common mistake. Choosing a retrieval design from a vendor demonstration and not from the mix of your own questions. Teams load tables into a vector index and are surprised when totals are wrong, or build an ontology for a contract library that needed text search.

How to check. Collect fifty likely questions and sort them into text, records and both. Then run five from each pile against the design you are considering, and note which pile produces wrong or unverifiable answers. Ask who checked each result.

What good looks like. The design follows the question mix. Record questions are answered by generated queries over live data, text questions by retrieval of passages, mixed questions by both, and someone owns the upkeep of the model.

01Foundation & governanceone set of numbers · lineage ·cataloguing · migration assurance02Operations & productionyield · asset performance ·condition-based maintenance03Quality & compliancegenealogy · long-horizon retrieval ·containment scoping04Commercial & revenuecost to serve · counterparty 360 ·revenue assurance05Finance & controlconsolidation · close acceleration ·working capital06Intelligence & AIself-service · insight agents ·document intelligence07Governed agentic operationsthreshold-bound action · approvalrouting · sealed record08Sensing & field operationsdetection with a coordinate · offlinecapture · twinsEight groups. Fifty-plus patterns. Read it as range rather than a menu:if a question can be answered from data you already hold, it can be built.
Exhibit 03What becomes answerable once the foundation exists.

Track B · part 4 of 7

B4. A system that never refuses tells you nothing when it answers

If every question receives a confident reply, the replies carry no information about which of them to trust.

Consider a colleague who answers every question with equal assurance, whether or not they know. You would soon stop relying on any of it. An AI system behaves the same way if it has no route to say that it cannot answer. Refusal is what gives an answer its meaning.

A refusal is not a failure state. It is a designed output with a reason attached, and there are several distinct reasons.

  • No evidence. Retrieval returned nothing relevant.
  • Conflicting evidence. Two sources disagree and the system cannot resolve them.
  • No entitlement. The person asking is not permitted to see the underlying data.
  • Out of scope. The question lies outside what the system was built and tested to handle.
  • Beyond authority. For an agent, the action exceeds the limit that a named person set.

Each reason needs its own wording, because a user who is told the system does not know will behave differently from one who is told they may not see the data. Each refusal should also be logged, since the log of refusals is a map of where your data and scope fall short.

For actions, the same idea applies as graded autonomy. EigenForge's Kyros platform, which is in production, grants autonomy by action class: some actions run automatically, some are prepared for a named person to approve, and some are advisory only. The principle works with any agent platform. Decide in advance which actions the system may take alone.

Now the difficult part. Setting the threshold at which a system refuses is an unsolved problem. Nobody has a method that works across domains, and anyone who says otherwise is selling something. The reasons are specific.

  • A model's stated confidence is poorly calibrated, so a high score does not reliably mean a correct answer.
  • Retrieval similarity scores are not comparable from one query to the next, so a fixed cut-off behaves differently across questions.
  • The cost of error varies. A wrong answer about a meeting room differs from a wrong answer about a contractual liability, yet a single threshold treats them alike.
  • Thresholds drift. A change of model, a data refresh or a new document type can shift the right setting without any warning.

Set it too high, and the system refuses useful questions and users drift back to spreadsheets. Set it too low, and wrong answers reach people who trust them. There is no setting that removes both problems.

The practical response is to treat the threshold as something you manage. Set it by question class, and not globally. Sample both the refusals and the answers on a regular schedule, and mark each as right or wrong. Track the refusal rate and the wrong-answer rate together, because either alone is misleading. Revisit the setting whenever the model, the data or the question mix changes.

In practice: Ask any vendor to show you, live, a question their system declines and a message explaining why. If every question in the demonstration is answered, ask what happens with a question about data the system does not have.

Next: B5, Evaluate on your own questions, because a leaderboard measures someone else's.

Put it to work

The common mistake. Setting one global confidence threshold during the demonstration and never revisiting it. After a model change or data refresh the right setting moves silently, and the system either refuses useful questions or answers badly with unchanged assurance.

How to check. Pull a sample of last month's refusals and a sample of answers. Mark each as right or wrong, and calculate the refusal rate and the wrong-answer rate separately for two question classes. Check whether each refusal states its reason.

What good looks like. Thresholds are set by question class, refusals carry distinct reasons, and both rates are tracked on a schedule. A change of model, data or question mix triggers a re-check, and the refusal log feeds data work.

FIVE WAYS TO SAY NO, EACH WITH ITS OWN WORDINGWHY IT REFUSESWHAT IT SAYSWHAT THE USER DOESNo evidenceNothing relevant wasretrieved.“I don’t have material that answersthis.”Asks elsewhere, ornarrows the question.Conflicting evidenceTwo sources disagree andcannot be reconciled.“Two records disagree; here areboth.”Investigates the source,not the system.No entitlementThe asker may not see theunderlying data.“You are not cleared for the recordsbehind this.”Requests access throughthe owner.Out of scopeOutside what the system wasbuilt and tested for.“That is not something I was builtto answer.”Routes to the right teamor tool.Beyond authorityAn agent action past thelimit a named person set.“This needs approval; I haveprepared it for review.”A person decides; thedecision is recorded.Each refusal is logged — the log is a map of where your data and scope fall short.
Exhibit 04Five designed refusals, each with its own wording and its own user behaviour.

Track B · part 5 of 7

B5. Evaluate on your own questions, because a leaderboard measures someone else's

A public benchmark tells you how a model performs on general tasks, and your data, vocabulary and permissions are not general.

Leaderboards are useful for a first shortlist of models. They cannot tell you whether a system will answer correctly about your customers, in your terminology, against your tables, for a user with restricted access. Only a test built from your own questions can do that.

A workable evaluation set has a few properties. It is drawn from real questions, taken from analysts' request backlogs, support tickets, meeting minutes and the queries people already run. It is large enough that one failure does not decide the verdict, which in practice means a few dozen to begin with, and it should grow. It includes the awkward cases deliberately.

  • Questions that require joining several systems.
  • Ambiguous questions, where the right behaviour is to ask for clarification.
  • Questions the system should refuse because the data does not exist.
  • Questions where two users should receive different answers because of entitlement.
  • Questions whose answer changed recently, to test freshness.

Each question needs an agreed correct answer, written by the person who owns the figure. This step is often skipped, and it is the most valuable one. It forces the organisation to decide what correct means, which is the same discipline as module A3.

Scoring should cover more than correctness. Record whether the answer was right, whether the evidence supports it, whether a refusal was appropriate, whether entitlement was respected, how long the response took and what it cost. A system that is slightly less accurate but never leaks restricted data may be the better choice. A single accuracy figure hides that trade.

Three practices protect the evaluation itself.

  1. Keep a held-out portion of the questions that nobody uses for tuning. Otherwise the system is optimised for the test.
  2. Have answers scored without knowing which system produced them, when comparing products.
  3. Rerun the whole set whenever the model, the prompt, the data or a definition changes. A system that passed in March has not necessarily passed in June.

The same set lets you compare vendors on equal terms, including products you have already bought. It also gives you a record of behaviour over time, which you will want when an auditor or a board member asks how you know the system still works.

Be cautious with automated scoring by another model. It is cheap and fast, and it has its own errors. Check a sample of its judgements by hand until you know how far to trust it.

In practice: Ask five people who use your data daily for the three questions they ask most often and the three they have given up asking. Write the correct answers with them. Those thirty questions are a usable first evaluation set, and they can be run against any product tomorrow.

Next: B6, Cost per answered question is the unit that makes systems comparable.

Put it to work

The common mistake. Writing the reference answers from the system's own output, so the evaluation confirms what the system already says. Or tuning against the same questions that are later used to score, so the result measures the tuning.

How to check. Look at your current test set. Ask who wrote the correct answer for each question, and whether it was written before or after anyone saw the system's output. Check that a held-out portion exists which nobody has tuned against.

What good looks like. A set of real questions with answers written by each figure's owner, including refusal and entitlement cases, a hidden portion, and a full rerun after any change to model, prompt, data or definition.

AN EVALUATION SET THAT CAN ACTUALLY DECIDEWHERE THE QUESTIONS COME FROM1Analysts' request backlogs2Support tickets3Meeting minutes4Queries people already run5The ones given up as unanswerableA few dozen to begin with.Large enough that one failure doesnot decide the verdict. It shouldgrow.THE AWKWARD CASES, ON PURPOSEJoins across several systemsAmbiguous — right behaviour is toclarifyRefusal — the data does not existEntitlement — two users, two answersFreshness — the answer changedrecentlySCORED, PER QUESTIONWas the answer right?Does the evidence support it?Was a refusal appropriate?Was entitlement respected?How long, and at what cost?A slightly less accurate systemthat never leaks restricted datamay be the better choice.Every question carries an agreed correct answer, written by the person who owns the figure.
Exhibit 05An evaluation set that can actually decide, and what is scored per question.

Track B · part 6 of 7

B6. Cost per answered question is the unit that makes systems comparable

A price per token or per seat says little until it is divided by the number of correct answers that people actually used.

AI systems are priced in many units: tokens, seats, queries, compute hours, platform licences. These cannot be compared directly, and a cheap unit can hide an expensive system. A single measure turns the offers into something you can set side by side.

Define it carefully. Cost per answered question is the total cost of running the system for a period, divided by the number of questions that received a correct answer which someone used. The denominator is important. Questions submitted is the wrong count, because it includes failures. Questions answered is better, but still includes wrong answers. Correct and used is the figure that matches the value you want.

The numerator has more parts than most budgets show.

  • Model usage. Input tokens, which include all the retrieved context sent with the question, and output tokens.
  • Retrieval and indexing. Embedding content, storing indexes and refreshing them as sources change.
  • Data platform. The queries run against source systems and the infrastructure that serves them.
  • Hosting. For local models, the hardware or reserved capacity, which is paid for whether it is busy or idle.
  • Evaluation and review. The testing described in module B5, and the human time spent checking answers.
  • Refusals and retries. A declined question still costs something, and so does every retry.

Several design choices move this number. Routing each question to the smallest model that can do the job, and using a larger model only when needed, usually changes cost more than any discount. Limiting the amount of retrieved context reduces input tokens. Caching repeated questions avoids paying twice. A budget that declines work at a set limit, instead of reporting an overrun afterwards, keeps a runaway process from becoming a surprise invoice.

Agents change the arithmetic. One user request may start many model calls, tool calls and retries, so the cost per request is no longer a fixed multiple of one call. Measure it per use case and per agent, and set a limit on each.

Local inference shifts the cost structure. A hosted model charges per token, so cost rises with use. A model you run yourself carries a largely fixed cost for capacity, so cost per question falls as utilisation rises, and rises when the hardware sits idle. Which is cheaper depends on your volume, and the comparison should use your own volumes and not a vendor's example.

Finally, record the cost next to the quality scores from B5. A cost figure without a quality figure rewards the system that answers badly and cheaply.

In practice: For one use case, collect last month's total spend across the six components above and divide it by the number of correct, used answers from a sample you have reviewed. If you cannot produce the denominator, that is the first gap to close. Without it, no price comparison is meaningful.

Next: B7. The token bill is five levers, and none of them is the price list.

Put it to work

The common mistake. Comparing suppliers on unit price and dividing spend by questions submitted. Retries, refusals, evaluation effort and idle hosting are left out, and wrong answers are counted as value, which rewards the cheapest system that answers badly.

How to check. For one use case, total last month's spend across model, retrieval, data platform, hosting, evaluation and retries. Divide by the correct answers that someone used, taken from a reviewed sample. If you cannot produce the denominator, start there.

What good looks like. Cost per correct, used answer is reported for each use case beside its quality score, with a limit that declines work. Routing and caching choices are judged by how they move that figure.

COST PER ANSWERED QUESTION, AS THE FOUNDATION IS REUSED0255075100RELATIVE COST PER QUESTIONEach question built from scratch1st2nd3rd4th5th6th7thQUESTIONS ANSWERED ON THE SAME GOVERNED FOUNDATIONREAD THIS AS A SHAPEThe axis is relative,not a currency. Thecurve is the argument:connection, definitionsand entitlement arepaid for once. Weestimate your ownnumbers from your ownrecords before anythingis committed.Illustrative. We publish no improvement percentages.
Exhibit 06Cost per answered question, as the foundation is reused.

Track B · part 7 of 7

B7. The token bill is five levers, and none of them is the price list

Most AI spend is self-inflicted: systems paying to re-read their own context, reasoning at length about lookups, and answering simple questions at frontier prices. The fixes are architectural, and they are known.

When the first production bill arrives, the instinct is to negotiate the price list. The price list is the least interesting part of it. Ungoverned agent workloads spend most of their money on things nobody chose: conversation history re-read on every turn, retrieved documents that are never used, tool schemas carried into every call, long reasoning about questions that deserved a lookup, and output tokens — priced at several times input — with no cap on length.

Five levers move most of this spend.

  • Route. Send each class of question to the smallest model that passes your own evaluation for that class. The price spread between tiers is large, and most enterprise traffic belongs in the cheap ones.
  • Cache. Stable prompt prefixes are charged at a fraction of the price by most providers; repeated questions can be answered from a semantic cache with no model call at all. Most enterprise workloads are repetitive, which is what makes this lever large.
  • Compress. Most tokens in a typical agent's input are low-signal. Compression before dispatch removes the majority of them without measurably hurting answers.
  • Cap. Output is the expensive direction. Cap answer length and reasoning depth per question class — short answers for short questions.
  • Budget. Every agent carries a spend counter enforced at the gateway, with anomaly alarms. This turns the other four from intentions into facts, and it is what catches a runaway loop at 3am.

Two details decide whether the levers actually work. Prefix caching pays only when the start of the prompt is byte-identical between calls, so prompt construction must be deterministic — stable parts first, variable parts last. And routing pays only when the evaluation behind it is honest: the cheapest model that passes a public benchmark is not the cheapest model that passes your questions on your data.

There is also a counter-movement worth knowing. As per-token prices fall, deliberately spending more tokens on the questions that deserve it — deeper reasoning, verification passes against the evidence — becomes affordable, and for high-error-cost questions it is the correct choice. The skill is not spending less. It is spending the right amount per question class, measured.

In practice: Take last month's bill and decompose it into the categories above: re-read context, reasoning, output, retries, and the questions themselves. Then apply one lever — routing is usually the largest — and measure cost per answered question before and after. If you cannot produce the denominator, that is the first gap to close.

Next: This is the last module in this track. Return to A1 and apply the questions in this series to your own estate.

Put it to work

The common mistake. Negotiating the model provider's price list while the system pays to re-read its own conversation history on every turn. The discount arrives; the bill does not move.

How to check. Decompose last month's spend into re-read context, reasoning, output, retries and the questions themselves. Then check whether prompt construction is deterministic — a timestamp or reordered section near the front of the prompt silently disables prefix caching.

What good looks like. A routing table derived from your own evaluation results, per-question-class caps on output and reasoning, and per-agent budgets enforced at the gateway — with cost per answered question reviewed monthly, because the models and the prices both move.

FIG. 1 ANATOMY OF AN UNGOVERNED AGENT BILL — SHARE OF SPEND, AND THE LEVER THAT MOVES IT38%Re-read context — history, retrieveddocuments, tool schemas22%Reasoning traces the task did not need18%Output — priced at four to six timesinput12%Retry and runaway loops nobody budgeted10%The question itselfTHE FIVE LEVERS THAT ACTUALLY MOVE ITROUTEeach call to the smallest model that passesyour own evaluationCACHEstable prefixes and repeated questions,semantically and exactlyCOMPRESScontext before dispatch — most input tokensare low-signalCAPoutput length and reasoning depth to whatthe task is worthBUDGETevery agent, enforced at the gateway, withanomaly alarmsUNIT THAT MATTERS: COST PER ANSWERED QUESTION —INCLUDING THE COST OF BEING WRONG
Exhibit 07Where the bill actually goes, and the five levers that move it.

Something here you disagree with?

These are written to be argued with. If a part of this is wrong, or missing the case you care about, tell us and we will fix it.

hello@eigenforgelabs.ai

Send opens your email client with the note already addressed to us — nothing is stored on this site, and the message goes from your own mailbox, so our reply lands in yours.