EigenForge AI Labs Talk to us
Making the case internally — track artwork

Learn · Track E

Making the case internally

The track nobody writes. Getting it approved, measuring it, and knowing when to stop.

Track E · part 1 of 5

A question the board can approve names a decision, an owner, a measure and a date

A board approves a decision it can picture being made differently. It cannot approve a capability.

Boards rarely refuse AI proposals because they distrust the technology. They defer them because the proposal describes an activity and not a result. "Deploy an assistant across finance" is an activity. A director reading it cannot tell what will be different in a year, who will notice, or how anyone will know that it worked.

A business question that a board can approve has four parts.

  • The decision. Which choice will be made differently, and by whom? A decision has a person attached and happens at a known point, such as a weekly credit review or a monthly purchasing run.
  • The measure. What will change in a number the board already recognises? Choose a measure that exists today, because E2 explains why you need its starting value.
  • The owner. Which executive answers for the result? This is the person whose budget or target moves, and not the sponsor who liked the demonstration.
  • The date. When will the board be told whether it worked, and what will it be shown?

Compare two versions of the same request. The first reads: "Use AI to improve collections." The second reads: "Can credit control identify, by the fifth working day of each month, which overdue accounts have open disputes, so that reminders are no longer sent to customers who are waiting on us?" The second names a decision (whom to chase), an owner (the head of credit control), a measure (reminders sent to disputed accounts) and a rhythm (monthly). It also reveals the data it needs: receivables and the dispute log, joined on a customer number. That is the problem described in A1, arriving early, and it is better to meet it on the first page than in the third month.

Three further habits strengthen the question. State what is out of scope, so that the approval cannot expand quietly. State what the current method costs in the time of named people, measured over a few weeks and not recalled from memory. And state the answer you would accept if it were negative. A question that cannot return "no" is a plan to buy something, and a board will sense that.

Language matters. Write the question in the vocabulary of the business and keep model names, vendor names and architecture off the first page. Technical choices belong in an appendix, where they can change without reopening the approval. A board asked to approve a technology has to judge it, and most boards cannot. A board asked to approve a decision can judge that with confidence, because judging decisions is its job.

Finally, check that the request is a single question. Proposals that bundle five use cases are common, because they feel efficient. They are hard to approve, since one weak case taints the rest and the board cannot approve part of it. Offer the strongest case first and list the others as options that depend on its result.

In practice: Rewrite your current proposal as one sentence containing a decision, an owner, a measure and a date. Read it to someone outside the project. If they cannot say what would be different in a year, the board will not be able to either.

Next: E2, The reference measure has to be captured before the project starts.

1Conversationthe thing that is actually inyour wayFREE2Written framingquestion, sponsor, referencemeasure, readinessYOURS TO KEEP3Proof of conceptyour data, your boundary, anhonest verdictSMALL, FIXED4Deliverybuild, validate, adoptSCOPED5Scalethe next question, at afraction of the costCOMPOUNDINGNo system replaced · no data migrated · no change to transaction processing · no multi-year programme to begin
Exhibit 01How an engagement starts. The first two steps cost nothing.

Track E · part 2 of 5

The reference measure has to be captured before the project starts, because afterwards it cannot be recovered

Without a starting value, every later claim of improvement is an opinion.

Most AI projects report success in the language of impressions. People say the work is quicker, the answers are better and the team is less stretched. These statements may be true. None can be shown, because nobody recorded how the work was done before. A reference measure is the starting value of the number you intend to change, taken before the project begins and by the method you will use afterwards.

Finding it is easier at the start than at any later point. Once a new tool is in use, people forget the old effort, the old records are archived and the team that did the work has moved on. A baseline rebuilt from memory will flatter the project, because people remember the old work as worse than it was.

The source is usually closer than expected.

What you want to changeWhere the reference often sitsCommon trap
Time to produce a report or answerTicket system, request backlog, timesheetsMeasuring elapsed days when the effort was hours
Error or rework rateAudit findings, credit notes, returned workCounting only the errors someone noticed
Volume handled per personWorkflow or case systemIgnoring differences in case complexity
Cost per transactionFinance ledger with allocationsAllocation rules that change between periods
Time to detect an issueIncident or complaint logRecording the date of discovery and not of occurrence

Four rules make a reference measure usable.

First, measure over a full cycle. Month-end, quarter-end and seasonal peaks behave differently from ordinary weeks, and a baseline taken in a quiet fortnight will make a later busy month look like a failure.

Second, measure the quality of the present method as well as its speed. The comparator for an AI system's accuracy is the person or spreadsheet that does the work today, and not perfection. That comparator makes mistakes. Recording how often lets you judge the new system fairly, and it stops the project being held to a standard nobody has ever met.

Third, write down how the figure was measured, so that it can be repeated identically: who counted, from which system, with which filters and over which dates. A different method afterwards produces a difference that comes from the method.

Fourth, ask the owner named in E1 to sign the figure. A baseline the owner has not accepted will be disputed at the review, which is the moment you most need it to stand.

Sometimes no reference exists, because nobody measures the work at all. In that case creating one is the first deliverable, and it is a legitimate result. It often shows that the process is smaller, or larger, than the proposal assumed. Either finding changes the case before any money is spent.

In practice: For the measure named in your question, find where the starting value is recorded today. If it is not recorded, ask the team to log it for one full cycle. Store the figure, the method and the date together, with the owner's agreement.

Next: E3, A proof of concept must leave behind a verdict, an evaluation set and a cost per answer.

Track E · part 3 of 5

A proof of concept must leave behind a verdict, an evaluation set and a cost per answer

A proof of concept that ends in a demonstration has produced a presentation. One that ends in evidence has produced a decision.

A proof of concept is a cheap way to learn, and it is cheap only if the learning is captured. Many are run, applauded and then repeated a year later by a different team, because the first left nothing behind except slides. Before you agree to run one, decide what it must produce. Six outputs separate a useful exercise from an expensive afternoon.

  1. A verdict against criteria written beforehand. The criteria come from the question in E1 and the reference in E2. They state what result means proceed, what means change course and what means stop.
  2. An evaluation set. Real questions with correct answers agreed by the owner of each figure, as B5 describes. This is the lasting asset, because it lets you test any product, including the next one, on equal terms.
  3. A measured cost per answered question. Use the definition in B6, with the review time included. A proof of concept that did not count cost has not tested whether the idea survives a budget.
  4. A list of what failed. Every wrong answer, every refusal and every question the system could not reach. This list is more informative than the successes and is usually the first thing to disappear from the readout.
  5. A test with a restricted user. Someone with narrower access than the administrator must have used the system, so that entitlement was exercised and not assumed.
  6. An operating sketch. Who would monitor quality, handle failures and own the spend if this went live. If nobody can say, the proof of concept has postponed a cost and not removed it.

Set the boundaries early. Fix a duration and hold to it, because an open-ended pilot gathers scope and loses its criteria. Use live data from at least one real source, even a narrow one, since a curated extract hides the join problems that decide the outcome. Keep the user group small but include people who did not ask for the project.

Settle ownership of the outputs in the agreement. The evaluation set, the logs and the written findings should belong to you and be usable without the supplier. If they sit in a vendor's tenancy, you cannot take them to a competitor, and the exercise has quietly become a lock-in.

Hold the readout with the person who signed the criteria in the room. Present the failures before the successes. Then record the decision, even when it is to stop. A recorded stop is a good outcome: it cost a small amount, it is documented, and it protects the organisation from running the same experiment again.

In practice: Take the proof of concept you are considering and check it against the six outputs. Any output it will not produce should be added to the scope or accepted as a known gap in writing. Agree the stop criteria before the first day of work.

Next: E4, Some internal objections are correct, and the proposal should concede them.

1Agree the questionprecise enough that twopeople recognise the sameanswer2Agree what success looks likein writing, and who decides —by name3Agree the boundaryread-only, nothing leaves, inthe topology you need4Build it against real datasomething a sceptical usercan try5An honest verdictincluding “this is not worthdoing”WHY MOST PROOFS OF CONCEPT FAILNo agreed success criterion, and no sponsor able to accept the result. Both are settled in the first hour.
Exhibit 02A proof of concept designed to produce a decision.

Track E · part 4 of 5

Some internal objections are correct, and the proposal should concede them

A case that treats every objection as resistance loses the objections that were right, and then loses the project.

Every AI proposal meets objections from finance, legal, security, operations and the people whose work will change. Sponsors often file them all under resistance to change. That is a mistake. Several objections are correct, and a proposal that concedes them in writing gains credibility for the rest.

The table sorts common objections into those that are usually right, those that are right under conditions, and those that usually are not.

ObjectionVerdictWhat settles it
"Our data is not ready"Right if the join keys between systems cannot be namedThe A1 test: name the systems, keys and owners
"Nobody is accountable when it is wrong"Right until an owner is namedA named executive and a written gate (C3)
"The time saved will not turn into savings"Right when savings are spread thinly across many peopleSay where the time goes: fewer hires, more cases, faster closing
"Security has not reviewed it"Right if entitlement is checked only at loginTest with two users of different rights (A4)
"We will be locked in"Right if evaluation sets and logs live with the supplierContract for ownership and export (E3)
"It is a solution looking for a problem"Right if the E1 question cannot be writtenWrite the question or stop
"The system will make mistakes"Wrong as stated, since the current method does tooCompare with the E2 reference measure
"Staff will resist"Usually about being left out, not about the toolInvolve the users from the first week
"Regulators will not allow it"Seldom true. Guidance asks for controls, not abstinenceMap the controls (C5) and ask counsel
"Wait until the technology settles"Partly right, but waiting has a costBuild the evaluation set now so any later switch is cheap

Three of the correct objections deserve emphasis. The first concerns savings. Time saved by many people, ten minutes each, rarely appears in a budget. If the case rests on capacity that nobody will redeploy, finance is right to discount it. Rewrite the case around a change the ledger will show, or around work that is currently not done at all.

The second concerns accountability. The objection is not that someone will be blamed. It is that today, in many proposals, nobody has been asked to answer for the outcome. Name the person before the board meets.

The third concerns data. Teams often reply with confidence that the data is fine. Test it. If a question needs two systems and nobody can state the join key, the objection is correct, and fixing identity is the project.

For the objections that are not right, respond with evidence and not with assurance. "The system will make mistakes" is answered by comparing its measured error rate with the reference rate. "Regulators will not allow it" is answered by the actual text and a conversation with counsel.

In practice: List every objection you have heard so far. Mark each as right, right under conditions, or not right, and write the evidence for the mark. Put the conceded ones, with your changes, on page one of the proposal.

Next: E5, Review the project after it ships, and treat stopping as a result.

Track E · part 5 of 5

Review the project on a fixed schedule after it ships, and write the stop criteria first

A system that is never reviewed drifts, and a project that cannot be stopped is not being managed.

Approval is the start of the evidence, and the board is owed the rest of it. Models change, data changes, questions change and users find ways to use the system that nobody predicted. A review after launch checks whether the case made in E1 still holds. Without one, the project continues on momentum.

Fix the review dates at approval, for example after the first full cycle, then at regular intervals. Put them in the board calendar. Reviews scheduled informally are the first to be postponed.

Assign the review to someone other than the sponsor. A sponsor reviewing their own project has every reason to find it successful. The reviewer should be able to read the numbers without needing the project team to explain them.

A review answers five questions.

  1. Is the measure moving? Compare the current value with the E2 reference, using the same method.
  2. Is it still accurate? Rerun the evaluation set from E3 and B5, since a model or data change may have shifted results since launch.
  3. What does it cost? Calculate cost per answered question as in B6 and compare with the figure at approval.
  4. Who is using it? Look at usage by team, and at the questions that fail or are refused. Declining use is information, and so is use that has moved to a task nobody planned.
  5. Is control intact? Check that every agent still has an owner, a scope and a budget, and that the gates in C3 have not been loosened without a decision.

Write the stop criteria before launch, when nobody is attached to the result. A criterion might be a measure that has not moved after an agreed number of review cycles, a wrong-answer rate above a stated limit, or a cost per answered question that exceeds the value of the work. The numbers are yours to set. What matters is that they were set in advance and in writing.

Stopping is a decision with options, and the choices are wider than switching off. You can stop completely, shrink to the use case that works, change the approach, or continue with conditions and a date. Sunk cost argues for continuing. The review should discuss only what happens next.

When the answer is to stop, close the work properly. Revoke agent identities and keys. Delete copies of data that existed only for the project. Keep the sealed records and the evaluation set, subject to retention rules, because an auditor may ask about decisions made while the system ran (see A5 and C4). Write a short note on what was learned and who should read it before the next proposal.

A project stopped on evidence has still done its job. The organisation now knows something it did not, at a cost it chose.

In practice: Before launch, put three dates and three stop criteria into the approval paper and name the reviewer. If a system is already live without these, set them this month.

Next: This completes the track. Return to E1 and use the same four parts to test the next proposal that reaches your desk.

Recovery & scrapMaterial recovered at the stage it was lostGross marginClaim & dispute evidenceContested rather than concededProvisionsContainment precisionScoped to units actually affectedCost of qualityInventory & receivablesWorking capital releasedFinancing costSpend consolidationNegotiation on the whole pictureCost of goods & servicesReport assembly effortHours returned to the businessOperating expenseTool consolidationSeveral licences become part of oneIT operating costNo published percentages. Each lever is estimated from your own operational and financial records before anything is committed.The two largest sources of value are the two we cannot quantify: decisions taken sooner, and decisions taken on one number.
Exhibit 03Seven levers, each landing in a measure finance already reports.

Something here you disagree with?

These are written to be argued with. If a part of this is wrong, or missing the case you care about, tell us and we will fix it.

hello@eigenforgelabs.ai

Send opens your email client with the note already addressed to us — nothing is stored on this site, and the message goes from your own mailbox, so our reply lands in yours.