Watch any agent demo this year and you will see the same trick: a capable model, a goal, and a loop that runs until the goal is met. It is genuinely impressive. It is also roughly fifteen per cent of what a production agent is.
The industry has a name for the other eighty-five per cent: the harness. The word comes from the safety equipment, and it is the right metaphor — the harness is not what climbs, it is what makes climbing survivable. Around every model that actually works for a living sits a structure of tools, memory, budgets, gates and records that decides what the model can touch, what it can spend, and what happens when it is wrong. The major framework releases of 2026 all converge on this shape. The model is rented. The harness is where your organisation lives.
Why the model is the interchangeable part
Model capabilities converge and prices fall; this year's frontier is next year's commodity tier. A system designed around one model's personality — its favourite prompt shapes, its particular failure modes, its context window — is a system that gets rebuilt every eighteen months.
A system designed around a harness absorbs model changes the way a well-run factory absorbs a new machine: the gauges, guards and shift procedures stay; only the machine is swapped. We route between models routinely, per question class, based on evaluation results. The reason we can do that without drama is that nothing about our governance lives in the model.
The technical detail: the layers that matter
A production harness has five layers, and each exists because of a failure that happens without it.
Tools with budgets. An agent acts through tools — read a ticket, query a ledger, post a draft. Each tool validates its inputs, makes its side effects idempotent, and carries a cost. "The agent did it seventeen times" is a budgeting failure before it is anything else.
Memory that is inspected. Agents need memory — what was tried, what worked, what was corrected. Memory that accumulates unsupervised becomes folklore. Ours is layered, reviewed, and never promoted into behaviour without a named person's approval.
A sandbox and a blast radius. Risky tools run restricted. Actions are allowlisted per environment. Anything irreversible — sending, deleting, charging — waits for a human. You design for the agent's worst day, not its average one.
Gates on the decision path. Authorisation lives at the single executor through which every effect passes, not in route middleware that a scheduled job can walk around. An action class is Auto, Approve or Advise, decided before anything is built, enforced the same way from every path.
Records and evaluation. Every lap of the loop leaves a signed record, and an evaluation loop measures the agent's work against outcomes, not vibes. An agent you cannot observe is an agent you cannot debug — and one you cannot reconstruct is one you cannot defend.
What this means for a buyer
When you evaluate an agentic offering, the demo will show you the model. Ask to see the harness instead.
Where do the tool budgets live, and what enforces them? Can you show me the record of one action, end to end? What happens at 3am when it loops? Which actions need my people's approval, and where is that enforced — in a guideline, or in the path the action has to travel? If I swap the model, what breaks?
A vendor with good answers to those questions has a product. A vendor with good answers about the model has a demo.
Talk to us
