Back to Blog
AI Strategy

What Did Your Agents Burn Last Month? Token Economics and the Retry Nobody Counted

40 to 60 percent of agent spend hides in the gap between the invoice and cost per successful outcome, and most of it is not the model. It is the retry.

Cost per token is a procurement metric. Cost per successful outcome is a management metric. Here is where the six sources of waste concentrate, and the six-week path to a real number.

September 4, 202615 min read
FinOpsToken EconomicsPrompt CachingCost AttributionAgent Runtime
AgentTrustOSCOST ENGINEERINGTOKEN ECONOMICSWhat Did Your AgentsBurn Last Month?The retry nobody countedWHERE THE WASTE HIDESFull-context retriesZero prompt cachingFrontier model, simple task20-35%of spend, in retries aloneSources: FinOps Foundation · OpenTelemetry gen_ai conventions · Epoch AI · AgentTrust OS 2026agent-trust.tech
Cost per token is a procurement metric. Cost per successful outcome is a management metric — and most organizations cannot yet produce the second one.
Cost Engineering · FinOps for AI · Agent Runtime
Key Facts
  • An agentic task — plan, retrieve, call tools, observe, re-plan, verify, produce output — consumes tens to hundreds of thousands of tokens, one to two orders of magnitude above the chatbot workload most budgets were sized against.
  • Analyses from Epoch AI and others put the decline in unit price for equivalent model capability at roughly an order of magnitude per year — yet total agent spend keeps climbing, a textbook Jevons effect combined with a genuine change in workload shape.
  • The FinOps Foundation has extended its scope to AI explicitly, with FinOps for AI guidance and the FOCUS billing specification adding coverage for model-inference costs, so spend can be normalized and allocated the way cloud spend has been for a decade.
  • OpenTelemetry's generative-AI semantic conventions define standard span attributes — model, input/output tokens — meaning cost attribution does not require a proprietary framework; a compliant agent can be priced within an existing observability pipeline in days.
  • Prompt caching charges roughly a tenth of the standard input rate on a cached prefix; batch processing runs at roughly half price for work that tolerates asynchronous completion.
TL;DR
  • Ask a CIO what one resolved claim or closed ticket cost, end to end, and the room goes quiet. That gap is where 40-60% of agent spend hides — and most of it is not the model.
  • Cost per token is a procurement metric. Cost per successful outcome is a management metric, and until you can produce the second one, every cost decision is a guess with a spreadsheet attached.
  • 20-35% of recoverable spend sits in retries that replay the whole conversation history on every attempt — an engineering defect, not a cost-versus-quality trade-off.
  • The cheapest token is the one you do not send twice. Fix caching and retries before touching model choice — no quality trade-off, and typically half the total saving.
  • A 45-65% reduction in cost per successful outcome is a realistic target inside one quarter, once attribution exists.
Keep reading → the six sources of waste, the six-week build sequence, and where AgentTrust OS fits.

Ask a CIO what their agents cost last month and you get an invoice. Ask what one resolved claim, one reconciled ledger, or one closed ticket cost, and the room goes quiet. That gap is where 40 to 60 percent of agent spend hides, and most of it is not the model. It is the retry.

Most organizations running agents today can answer exactly one financial question: what the provider charged last month. It arrives as a handful of line items — a model, a token count, a total — and it cannot answer which of fourteen agents consumed the majority of spend, what one resolved case cost end to end, or how much of the bill went on runs that ultimately failed. The absence of these answers has a predictable consequence: finance sees a line growing 20-30% a month with no unit denominator, and imposes a cap. The cap lands on the agents that are working, because they are the ones with volume.

The distinction worth internalizing

Cost per token is a procurement metric. Cost per successful outcome is a management metric. Until you can produce the second one, per agent and per task type, every cost decision is a guess with a spreadsheet attached.

Where the money actually goes

Token spend is not a model-selection problem. It is an accounting problem followed by a control-loop problem, and it is worth attacking in that order.

01
Attribute before you optimize
Two weeks of instrumentation reliably changes which optimization the team was about to spend a quarter on. The biggest line is almost never the one anyone predicted.
02
Measure cost per outcome, not per call
Failed runs consume tokens and deliver nothing. Any metric that excludes them flatters the system and hides the retry problem entirely.
03
Enforce budgets in the runtime
A monthly cost review cannot stop a runaway loop that burns a month's budget in an afternoon. The ceiling has to be a per-run token limit and a guard that can refuse a call.
04
Fix retries before you change models
Model swaps are visible, political, and risky. Retry discipline is invisible, uncontroversial, and usually larger.

The five numbers that matter

MetricDefinitionWhy it is the one to watch
Cost per successful outcomeTotal spend on a task type ÷ successful completionsThe only figure comparable to the manual process it replaced
Waste ratioSpend on failed, abandoned, and superseded runs ÷ total spendWhere the recoverable money is — above 20% is common and fixable
Retry amplificationMean tokens per completed run ÷ tokens of a first-attempt runExposes full-context replay. Above 1.6 means the retry path is naive
Cache hit rateCached input tokens ÷ total input tokensFor a stable agent this should exceed 70%. Most start near zero
Frontier shareSpend on the largest model ÷ total, split by task complexityReveals frontier models doing classification work

The six places waste concentrates

The shares below are indicative ranges rather than a fixed distribution, but the ordering has been remarkably stable across deployments.

20–35%

Retry with full-context replayThe same history re-sent on every attempt. Signal: retry amplification above 1.6. Fix: delta repair and a circuit breaker.

15–30%

No prompt cachingSystem prompt and tool definitions re-read every call. Signal: cache hit rate near zero. Fix: stable prefix and cache breakpoints.

10–25%

Over-modelling simple stepsFrontier model doing classification and formatting. Signal: frontier share high on low-complexity tasks. Fix: model cascade with escalation.

8–20%

Context bloatWhole documents pasted, history never pruned. Signal: input tokens rising with conversation length. Fix: retrieve spans, summarize, prune.

5–15%

Unbounded loops and fan-outNo step limit, multi-agent broadcast with no budget. Signal: a long tail of runs at the token ceiling. Fix: step budget and per-run ceiling.

3–10%

Interactive pricing on batchable workEvals, back-office runs, overnight processing. Signal: synchronous calls outside business hours. Fix: batch endpoint, roughly half price.

The first two rows require no quality trade-off at all — they are engineering defects rather than cost-and-quality decisions, which is why they come before any model-downgrade conversation.

The anatomy of the largest line

Retrying is correct behavior. The problem is the standard implementation: when a tool call fails validation, most agent frameworks append the error to the conversation and re-send the whole thing. Context grows with each attempt, and because input tokens dominate agentic workloads, cost grows with it.

Retry Anatomy: Full Replay vs. Delta RepairPATTERN A — naive retry, the whole conversation goes back every timeattempt 1 · 12kattempt 2 · 26kattempt 3 · 54k92k tokens, one outcome, still unverifiedPATTERN B — delta repair, only the typed error and failing fragment go backattempt 1 · 12kattempt 2 · 13.2kbreaker → escalate25.2k tokens — 3.6× less, bounded worst caseThe saving does not come from retrying less. It comes from changing what gets re-sent.
Figure 1: Retry anatomy. Delta repair with a circuit breaker also produces a bounded worst case per run, which is what makes capacity planning possible.
Retry causeWhat usually happensWhat works better
Transient infrastructureImmediate retry, same payloadExponential backoff with jitter, idempotency key, no context change
Schema/format violationAppend error, resend everythingConstrained decoding first; if it fails, send the typed error and failing fragment only
Business rule failureResend with "you made a mistake"Return the specific rule violated and the field
Judge/guardrail rejectionRegenerate from scratchCap regeneration at one attempt; a second rejection is a routing decision
Planner thrashNothing, it runs until the timeoutProgress detection, a hard step budget, abandon then escalate

The four structural levers, in order

1

Cache the stable prefix. Everything invariant — system instructions, tool definitions, policy text, few-shot examples — sits at the front and never changes within a session. Usually a one-day change, the largest single return.

2

Fix the retry path. Typed errors, delta repair, circuit breakers, step budgets. Two to three weeks, no quality trade-off.

3

Route by complexity. A cascade — small model attempts, escalates on low confidence — rather than one model for every step. Needs eval coverage to do safely.

4

Batch what does not need to be interactive. Evals, overnight processing, bulk document work. Pure scheduling change, roughly half price.

Six weeks to a real number

Week 1

Span schema agreed and emitted from one agent. Rate card loaded, versioned, effective-dated.

How you know it worked: any single run can be priced end to end.
Week 2

All agents instrumented. Outcome join on run identifier. First burn report, unedited.

How you know it worked: cost per successful outcome exists per task type.
Week 3

Waste analysis: retry amplification, cache hit rate, frontier share, the expensive tail.

How you know it worked: a ranked list of waste sources with a value on each.
Week 4

Prompt restructured for caching. Per-run token ceiling and step budget enforced in the runtime.

How you know it worked: cache hit rate above 60%, no run exceeds the ceiling.
Week 5

Retry path rebuilt: typed errors, delta repair, circuit breaker, escalation instead of a third attempt.

How you know it worked: retry amplification below 1.3.
Week 6

Routing cascade on the highest-volume simple task, gated by the eval suite. Batch migration for offline work.

How you know it worked: no quality regression on the eval gate.

What it is worth

A worked example on an illustrative mid-size deployment, 250,000 runs a month. The rate card is representative rather than any provider's current pricing — the ratios between the lines are far more stable than the absolute figures.

LineBeforeChange madeMonthly saving
Input tokens billed at full rate92%Stable-prefix caching → 28%$34,000
Retry amplification1.9×Typed errors, circuit breaker → 1.2×$21,000
Frontier model share of calls100%Cascade on 3 simple task types → 54%$16,000
Offline work on interactive pricing18%Batch endpoint → 2%$7,000
Runs abandoned at timeout4.1%Progress detection, escalation → 0.6%$5,000
Cost per successful outcome$0.61—$0.24 · 61% reduction
AgentTrust OS: Where the Burn Actually Shows Up

The six sources of waste above are not a separate cost project. AgentTrust OS's runtime intercepts the exact decision points where the burn concentrates — before it becomes an invoice line nobody can attribute.

Trust Certify
Catches the expensive pattern before it reaches production traffic. Running the same task thousands of times through the Reliability Engine surfaces retry loops and prompt bloat while the cost is still a test bill, not a production one.
Trust Runtime
Is a deterministic-first gate: nine rule-based checks resolve most decisions in under 20ms, and an LLM judge is invoked only where genuinely needed — removing the retry-on-a-slow-expensive-model path for every decision a cheap check can already settle.
Trust Audit
Is where cost per successful outcome stops being a spreadsheet exercise. Every decision carries its confidence score and routing outcome, so the monthly burn report is a query against the audit trail, not a reconciliation project someone runs once.

Frequently Asked Questions

Model downgrade is the lever executives reach for first and the one most likely to backfire. A small model that fails and escalates has cost you both calls plus a latency penalty, and cascades only pay when the small model's success rate on the routed slice is genuinely high — which you can only know from eval data. Fix caching and retries first: they require no quality trade-off at all and typically deliver half the total saving before a single model-choice decision gets made.
No, and the difference is where most of the waste hides. Failed runs consume tokens and deliver nothing — a metric that excludes them flatters the system and hides the retry problem entirely. Cost per call also cannot tell you which of your fourteen agents actually consumed the majority of spend, or what one resolved case cost end to end. You need the outcome join on run_id to get the number that is actually comparable to the manual process the agent replaced.
If your agents already emit OpenTelemetry gen_ai semantic convention spans, days rather than months — you build a cost ledger over your existing observability pipeline. The two-week instrumentation phase in this post's timeline is deliberately front-loaded, because you cannot recover token data you did not capture; retrofitting is impossible once the tokens are gone.
It is the highest-return single change available precisely because it is mechanical, not a quality trade-off. Re-reading a cached prefix costs roughly a tenth of the standard input rate. For agents whose system prompt, tool definitions, and policy text run to tens of thousands of tokens and are identical on every call — which is most agents — restructuring the prompt so invariant content sits at the front is usually a one-day change with the largest single return on the list.
It reduces the specific waste this post is about, and it does so by design rather than by accident: Trust Runtime resolves most decisions with nine deterministic, rule-based checks in under 20 milliseconds, invoking an LLM judge only for the dimensions that genuinely need one. That ordering removes exactly the retry-on-a-slow-expensive-model path this post identifies as the largest recoverable line. What it does not do is optimize your agent's own prompt structure or model choice — that work is still yours, and the burn report is what tells you where to spend it.
Ready to See Your Real Number?

No AI Agent Enters Production Without AgentTrust

A deterministic-first runtime that resolves most decisions in under 20ms, and an audit trail that turns cost per outcome into a query, not a project.

Start Free →

More from the blog

AI ComplianceJuly 22, 2026AI ComplianceJuly 22, 2026AI GovernanceJuly 22, 2026AI GovernanceJuly 22, 2026AI ArchitectureJuly 22, 2026AI SecurityJuly 16, 2026MLOpsJuly 10, 2026AI ImplementationJuly 8, 2026AI Agent ArchitectureJuly 5, 2026AI Agent ArchitectureJune 30, 2026EngineeringJune 23, 2026AI Agent ArchitectureJuly 2, 2026AI Agent ArchitectureJuly 1, 2026IntegrationsJuly 1, 2026IntegrationsJuly 1, 2026IntegrationsJuly 2, 2026AI SecurityJuly 3, 2026AI SecurityJuly 1, 2026AI ComplianceJuly 2, 2026AI ComplianceJuly 3, 2026AI StrategyJuly 2, 2026AI StrategyJuly 3, 2026AI StrategyJuly 3, 2026AI StrategyJuly 3, 2026AI GovernanceJuly 28, 2026Healthcare AIJuly 28, 2026ArchitectureJuly 29, 2026ArchitectureJuly 29, 2026AI StrategyJuly 29, 2026AI StrategyJuly 30, 2026AI SecurityJuly 30, 2026AI ComplianceJuly 30, 2026EngineeringJuly 30, 2026AI GovernanceAugust 4, 2026EngineeringAugust 4, 2026EngineeringAugust 4, 2026Agent SecurityAugust 19, 2026AI ComplianceSeptember 4, 2026EngineeringSeptember 4, 2026AI GovernanceSeptember 4, 2026