An agent whose work a human has to fully redo in order to trust it has not saved anything. Here is why verification is a design choice, not a review process.
14 minutes is a typical human time to verify one agent output that took the agent four seconds to produce. Here is the verification hierarchy and trust ladder that gets it to 2.6.
Ask how long it takes a person to check one agent output. If nobody has ever timed it, that number is almost certainly where the business case went. An agent whose work a human has to fully redo in order to trust it has not saved anything. It has moved the cost from production to verification and added a licence fee.
An agent is built in six to eight weeks. It works. It goes live with a sensible control: a human reviews every output before it takes effect. Everyone agrees that is prudent for the first few months. Twelve months later the review step is still there. It has acquired a team, a queue, a service-level target, and a headcount line. The agent generates a draft in four seconds and a person spends fourteen minutes checking it, because to check it properly they have to open the same source documents, re-read the same policy, and re-derive the same conclusion. The programme reports eleven thousand cases automated. The finance view shows a cost line that went up.
The most under-modelled cost in enterprise AI. It never appears in the business case, because the case compared the agent's cost to the manual process and quietly forgot that the manual process is still running, just relabelled as review. It is permanent by default — nobody is ever rewarded for removing a control that has not yet failed. It degrades — humans are poor at sustained monitoring of a mostly correct process, so review quality falls as accuracy rises, exactly backwards from what the control assumes. And it is invisible in the AI budget, because review headcount sits in the operating unit, not the technology programme.
The most striking recent study is METR's randomized controlled trial with experienced open-source developers working in their own repositories. Developers using AI tooling took roughly 19 percent longer to complete their tasks, while believing they had been about 20 percent faster. The perception gap is the finding that matters — generation felt fast, and the review, correction, and integration steps that consumed the time were not what participants were counting.
The 2024 DORA Accelerate report pointed the same way at an organizational level, associating increases in AI adoption with small decreases in delivery throughput and larger decreases in delivery stability, even as developers reported feeling more productive. Neither result says AI does not work. Both say the same thing: the cost moved downstream and the measurement did not follow it.
Lisanne Bainbridge's 1983 paper "Ironies of Automation" identified this exact trap in industrial control: automating the routine part of a task leaves the human with the residual task of monitoring — the task humans perform worst — while eroding the practice that made them competent to intervene. Every AI review queue is a re-run of that experiment.
The useful idea from computer science is verification asymmetry: checking a proposed answer is frequently far cheaper than producing one, but only when the answer arrives in a form that supports checking. A sorted list is trivial to verify. A three-paragraph narrative conclusion with no sources costs as much to verify as it did to produce. Same accuracy, wildly different economics — and almost nobody pulls this lever. Most agent outputs are optimized for reading when they should be optimized for checking.
Minutes of effort per completed case, illustrative, from a claims-adjudication workflow.
Net value per item = C(manual) − [ C(agent) + C(verify) + p(escape) × C(failure) ]. C(verify) dominates — in the middle lane above it is more than 80% of post-automation cost, yet it receives almost none of the engineering attention. C(agent) is the smallest term and gets the most airtime.
An output with these five properties can usually be checked in one to three minutes instead of ten to twenty, with higher confidence rather than lower.
Design rule: every check belongs at the lowest tier capable of performing it.
The failure state: most deployments operate almost entirely in Tier 6. Every item moved down a tier is a permanent, compounding saving.
Once the automated tiers exist and you have a measured defect rate, blanket human review becomes statistically indefensible as well as expensive. Manufacturing settled this decades ago with acceptance sampling (ISO 2859-1), and the logic transfers directly: define the acceptable quality level per risk tier, sample stratified — never uniformly, weighted toward high-value items and low-confidence outputs — tighten automatically on evidence, and publish the operating characteristic so everyone can see what defect rate the current plan would detect.
| Rung | Human review | Evidence needed to climb | What sends it back down |
|---|---|---|---|
| L0 · Draft only | Human authors every item | Starting position | Not applicable |
| L1 · Approve every action | Agent acts only after sign-off | Eval gate green, verifiable output shipped | Any critical finding |
| L2 · Approve high-risk only | Low-tier actions execute themselves | 1,000 items at or below target defect rate | Defect rate above the agreed level |
| L3 · Sampled review | 5-20%, stratified, with tightening rules | Two consecutive clean sampling periods | Two defects in one sample |
| L4 · Exception only | Humans see flagged and abstained items, under 5% | Sustained performance plus a rollback drill | Any escaped defect at severity high |
The ladder only works if promotion criteria are published in advance, demotion is automatic rather than discretionary, and the current rung is visible to the business owner.
Measure the tax. Time and motion on the existing review step, per item, per reviewer.
How you know it worked: a number everyone agrees is real.Decompose the output. Restructure the response into atomic, individually checkable claims with per-claim confidence.
How you know it worked: a reviewer can disagree with exactly one claim.Wire the sources. Deep links that resolve to the exact clause or record.
How you know it worked: zero unresolvable citations in production.Tier 1 and 2 checks. Deterministic recomputation of everything derivable, policy rules, reconciliation.
How you know it worked: most of the checkable surface auto-cleared.Independent verifier. Tier 3 verifier from a different model family, scoring claim support.
How you know it worked: verifier calibrated against human labels.Rebuild the console. Collapse verified claims, sources in place, amend flow, automatic capture of corrections as eval cases.
How you know it worked: median review time cut by more than half.Sampling plan. Risk-tiered acceptance quality levels, stratified sampling, automatic tightening rules.
How you know it worked: second line signs the sampling plan.Publish the ladder. Autonomy rungs, promotion criteria, automatic demotion triggers, a rollback drill.
How you know it worked: a drill demotes an agent successfully.Using the illustrative figures above, at 120,000 cases a year and a fully loaded reviewer cost of $60/hour.
| Operating model | Human min/case | Annual human cost | Versus manual |
|---|---|---|---|
| Manual baseline | 22.0 | $2.64M | none |
| Agent, blanket review | 17.0 | $2.04M | 23% saved |
| Verification-first, sampled review | 2.7 | $0.32M | 88% saved |
The comparison that matters is the last line, not the first. Both agent rows use the same agent. What differs is how the work is presented for checking — a difference worth roughly three times what the automation itself delivered.
The trust ladder above is only real if it is enforced at runtime, every time, rather than designed on a whiteboard. AgentTrust OS routes decisions to the rung they have earned and generates the evidence that lets a decision type climb.
Runtime-enforced trust tiers, an audit trail that makes a decision cheap to check, and autonomy that is earned rather than assumed.
Start Free →