EigenForge AI Labs Talk to us

Blog

The five levers on an AI bill

Enterprise AI bills are not driven by one expensive model call. They are driven by systems that pay to re-read the same context all day. Five levers move most of the spend, and none of them is negotiating the price list.

← All articles

2026-10-05 · EigenForge AI Labs

There is a conversation happening in most enterprises right now that goes like this: the pilot was cheap, the rollout is not, and someone is asking whether the model provider can sharpen its pencil. The provider's price list is the least interesting part of the bill.

The bill is mostly self-inflicted. Agent systems pay to re-read their own conversation history on every turn. They retrieve documents, stuff them into context, and use a tenth of them. They carry tool schemas the size of short stories into every call. They reason at length about lookups, and they answer simple questions at frontier prices because nobody told them not to. Output tokens cost several times more than input tokens, and nobody capped the output.

None of this is a criticism of the models. It is a description of systems assembled quickly by people who were, correctly, more interested in whether the thing worked than in what it cost. Now it works. Time to fix the bill.

The five levers, in order of usual return

Route. Different questions deserve different models, and the price spread between the smallest competent model and the frontier is ten to fifty times. Routing each call to the smallest model that passes your own evaluation for that class of question is the single largest lever in most estates — industry measurements put typical savings between a fifth and four-fifths of spend, depending on the workload mix. The routing rule is data, not vibes: it comes from your evaluation results, and it changes as models improve.

Cache. Much of what an agent reads, it has read before. Stable prompt prefixes can be cached by the provider at a fraction of the price. Repeated questions can be answered from a semantic cache without a model call at all. For workloads with repetitive patterns — and most enterprise workloads are repetitive — this is the difference between paying for a question once and paying for it every day.

Compress. Most tokens in a typical agent's input are low-signal: boilerplate, history that no longer matters, retrieved passages that turned out to be irrelevant. Compression before dispatch routinely removes most of them without measurably hurting the answer. It is unglamorous and it works.

Cap. Output is the expensive direction. Capping answer length and reasoning depth to what the task is worth — short answers for short questions — costs nothing in quality when it is set per question class rather than globally.

Budget. Every agent gets a spend counter, enforced at the gateway, with anomaly alarms. This is the lever that turns the other four from good intentions into facts — and it is the one that catches the runaway loop at 3am that would otherwise run until somebody's card statement does.

The technical detail

Two mechanisms carry most of the weight, and both are worth understanding one level down.

Prefix caching only pays when the beginning of the prompt is byte-identical between calls, which means prompt construction has to be deterministic: the stable parts first, the variable parts last, no timestamps or random ordering creeping into the front of the context. Teams that sprinkle volatile content into their system prompts pay full price and never learn why.

Routing only pays when the evaluation behind it is honest. The cheapest model that passes a benchmark is not the cheapest model that passes your questions on your data. We measure cost per answered question per question class, including the cost of being wrong, from a per-call ledger that records model, tokens, latency and a difficulty score for every request. The routing table falls out of that ledger. Without the ledger, routing is a guess with good manners.

FIG. 1 ANATOMY OF AN UNGOVERNED AGENT BILL — SHARE OF SPEND, AND THE LEVER THAT MOVES IT38%Re-read context — history, retrieveddocuments, tool schemas22%Reasoning traces the task did not need18%Output — priced at four to six timesinput12%Retry and runaway loops nobody budgeted10%The question itselfTHE FIVE LEVERS THAT ACTUALLY MOVE ITROUTEeach call to the smallest model that passesyour own evaluationCACHEstable prefixes and repeated questions,semantically and exactlyCOMPRESScontext before dispatch — most input tokensare low-signalCAPoutput length and reasoning depth to whatthe task is worthBUDGETevery agent, enforced at the gateway, withanomaly alarmsUNIT THAT MATTERS: COST PER ANSWERED QUESTION —INCLUDING THE COST OF BEING WRONG
Exhibit 01Where the money actually goes, and the five levers that move it.

The uncomfortable corollary

Per-token prices are falling fast — and total enterprise AI spend is rising anyway, because cheaper tokens make longer reasoning affordable and longer reasoning is where the value is. Analysts have started warning procurement teams not to confuse the deflation of commodity tokens with the democratisation of frontier reasoning. The organisations that come out ahead are not the ones paying the least per token. They are the ones whose cost per correct, defensible answer keeps falling.

That number is measurable. Ask for it.

costtokensfinops

Disagree with this?

These are written to be argued with. If you think this is wrong, we would rather hear it than not.

hello@eigenforgelabs.ai

Send opens your email client with the note already addressed to us — nothing is stored on this site, and the message goes from your own mailbox, so our reply lands in yours.