AI Agent Cost Per Task: Why Your Bill Is a Reliability Problem
Token cost grows with the square of the steps, the chance of needing another attempt grows exponentially with them, and the human who cleans up afterwards costs more than both. Here is the model we use to decide whether an agent is viable, with the arithmetic shown.

There is a benchmark result that explains most of what goes wrong when an AI agent moves from a demo to a production queue, and it has nothing to do with which model you picked.
In a July 2026 analysis of long-horizon evaluations, Peng and colleagues report figures from SWE-Milestone in which an agent scores roughly 80% or higher when each milestone is attempted independently from a clean snapshot, and 38.03% when the same milestones are run continuously as one task. On the scikit-learn subset, one configuration falls from 93.2% independently to 21.1% continuously. Same model. Same work. Same tools. The only thing that changed was whether the agent had to carry its own history forward.
That gap is the entire subject of this article, because it is simultaneously a reliability number and a cost number. An agent that fails is an agent that retries, and a retry does not cost you one step — it costs you the whole conversation again, at a context length that has been growing the entire time. The bill and the failure rate are the same phenomenon measured in different units.
Most writing about agent cost stops at "agents use more tokens than chatbots" and offers a monthly range. That is not actionable. What follows is the model we use to decide whether an agent is viable, where the money actually goes, and which architectural changes move the number — with the arithmetic shown so you can substitute your own figures.
The only cost number that means anything
Cost per token is a procurement metric. Cost per API call is a billing artifact. The number that decides whether an agent is worth running is fully loaded cost per resolved task: everything you spend on a unit of work, divided by the units of work that actually came out finished and correct.
Written out, for a task entering the system:
| Term | What it is | Usually ignored because |
|---|---|---|
| Attempt cost | Input + output tokens for one complete run, including tool-result tokens | Teams estimate it per step and forget the context re-send |
| Expected attempts | How many runs it takes before one succeeds | Pilots run on clean inputs where it is close to 1 |
| Escalation cost | What a human costs when the agent gives up or is wrong | It lands in a different budget line from the API bill |
| Verification cost | Judge models, test runs, deterministic checks | It is small per call and invisible in aggregate |
| Rework cost | Cleaning up confidently wrong output that shipped | Nobody attributes it back to the agent |
The structure matters more than the precision. Once you write it this way, something becomes obvious that the per-token framing hides: the human escalation term is usually the largest one, and it is controlled entirely by the success rate. Cutting your model price in half does nothing to it. Raising your resolution rate by ten points cuts it by ten points.
Why the token bill grows faster than the work
An agent loop is stateless at the API layer. Every step re-sends everything that came before it. That single property is responsible for most of the gap between what teams estimate and what they are billed.
At step i, you send the system prompt and tool definitions (call it S), plus everything accumulated from the previous steps. If each step adds roughly a tokens of model output and tool results, then total input tokens across n steps is:
total input = n·S + a · n(n-1)/2
total output = n·o
The second term is quadratic. Doubling the number of steps does not double the input bill; it roughly quadruples it. With a 4,000-token system-and-tools prefix, 1,500 tokens added per step, and 400 output tokens per step:
At Claude Sonnet 5 list prices (2 USD per million input tokens, 10 USD per million output), the ten-step run costs about 26 cents and the forty-step run about 2.82 USD. Four times the steps, eleven times the cost. Nothing went wrong in either run — that is the price of success.
Why success falls off a cliff, and why that is not mysterious
The second force is reliability, and the honest version of this story is less dramatic than the usual telling but far more useful.
If each step of a task succeeds independently with probability p, then an n-step task succeeds with probability pn. That is it. Peng and colleagues call this the "auditable null model" and make the point that most of the alarming decline people attribute to a special long-horizon weakness is what you would predict anyway from ordinary per-step error compounding. Four repairs at 80% each predict 41% overall. A 12-step task at 95% per step predicts 54%.
Two independent lines of work point the same way. Toby Ord's half-life analysis (arXiv, May 2025) models agents as having a constant hazard of failure per unit of task length, which produces exponentially declining success as tasks get longer and gives each agent a characteristic half-life. And METR's time-horizon work reports that the task length a model can complete at 80% success is four to six times shorter than the length it manages at 50%. The headline horizon numbers you see quoted are the 50% figures. Production rarely accepts a coin flip.
None of this says agents are bad. It says horizon is the dominant variable, and horizon is something you control in architecture, not something you buy from a vendor.
Putting the two together
Now combine them, because this is where the decision actually gets made. Token cost per attempt grows roughly with the square of the steps. The probability of needing another attempt grows exponentially with the steps. They multiply.
For a task entering the system, with per-attempt cost C, success probability p, a retry cap of k attempts, and a human escalation cost H:
expected attempts = (1 - (1-p)^k) / p
eventual resolution = 1 - (1-p)^k
cost per task = C · expected attempts + (1-p)^k · H
Take the ten-step agent above: C = 0.26 USD, and suppose it completes 70% of tasks unaided, retries once on failure, and that a human picking up the exception costs 6 USD fully loaded. Substitute your own H — it is the number most worth measuring honestly.
| One 10-step agent | Two checkpointed 5-step stages | |
|---|---|---|
| Billed tokens per clean run | 107,500 | 70,000 |
| Cost per attempt | $0.26 | $0.09 per stage |
| Success without intervention | 70% | 84% per stage |
| Resolution after one retry | 91% | 94.7% overall |
| Model spend per task | $0.34 | $0.21 |
| Escalation spend per task | $0.54 | $0.32 |
| Cost per task | $0.88 | $0.53 |
The right-hand column is the same agent doing the same work, split into two stages with a validated handoff between them, each stage retried on its own. Per-step reliability is held constant at the rate implied by the left column — nothing got smarter. The improvement comes from two places at once: each attempt is shorter, so the quadratic term is smaller, and each retry replays one stage instead of the whole task.
One caveat worth stating plainly, because it is the kind of thing that gets dropped when a model like this is repeated: independence between steps is an assumption, not a law. Real agents fail in correlated ways — a bad plan at step two poisons everything after it. Peng and colleagues argue this null model should be the baseline you measure against, not the answer. Use it to size the problem, then measure your own per-stage rates.
What actually moves the number
Ranked by the size of the effect we see in practice, not by how often they are recommended.
1. Shorten the horizon
Everything above is an argument for this. Decompose the task into stages with a validated handoff, persist the state between them, and retry at the stage level. A stage boundary is also the natural place to put a deterministic check, a human approval, or a cheaper model.
The useful discipline is a horizon budget: decide up front the maximum number of model turns any single uninterrupted trajectory is allowed, and treat exceeding it as a design failure rather than a runtime condition. In most production systems we would be uncomfortable above roughly ten turns between checkpoints, and happier at five.
2. Cache the stable prefix
The system prompt and tool definitions are re-sent on every single step and they never change. That is exactly what prompt caching exists for, and the published multipliers make the arithmetic easy. On the Claude API, cache writes cost 1.25x the base input price for the five-minute TTL and 2x for the one-hour TTL, while cache reads cost 0.1x on most models. Across a ten-step loop, a 4,000-token prefix goes from 40,000 billed tokens to the equivalent of about 8,600 — a 78% reduction on that portion.
Two things ruin it, and both are easy to do by accident. Caching is a prefix match, so any byte that changes anywhere in the prefix invalidates everything after it — a timestamp in the system prompt, a reordered JSON key, a tool list built from an unsorted map. And there is a minimum cacheable length, between 512 and 4,096 tokens depending on the model; below it nothing is cached and no error is returned. Check usage.cache_read_input_tokens on real traffic rather than assuming.
3. Stop carrying dead context
Most of what accumulates in an agent loop is tool output that mattered for one step and is inert afterwards: a 6,000-token API response the model extracted one field from, a directory listing, a failed attempt. Clearing or summarising those is the difference between the quadratic term compounding and the quadratic term being bounded. Both major providers now ship first-party mechanisms for this; the design question is which results are safe to drop, and that is answered by your task, not by the vendor.
4. Route by difficulty, not by default
Most steps in a typical agent trajectory are mechanical: parse this, pick the next tool, format that. They do not need the frontier model. Routing those to a cheaper tier and reserving the expensive model for planning and ambiguity is a large, boring win.
Two honest caveats. Caches are model-scoped, so a cascade forfeits cache reuse across its tiers — measure the combination, not each lever alone. And before building a router, try the simpler thing: the capable model at a lower reasoning effort, on the same tasks. It often beats a cheaper model at high effort, and it keeps you on one cache namespace.
5. Cap retries, and make the cap intelligent
Unbounded retry is how pilots become incidents. But a flat cap wastes money too: the second identical attempt on a task that failed for a structural reason will fail identically. A retry is only worth paying for when something about the attempt changes — a different model, a repaired input, a reduced scope, or a tool that was unavailable the first time. If your retry policy cannot name what will be different, it should be an escalation instead.
6. Verifiers: useful, but smaller than advertised
This is where we would push back on the prevailing advice. In a July 2026 decomposition of a production enterprise agent, Dastidar and the Leni team separated the contributions of their architecture on SpreadsheetBench: total uplift of 11 percentage points, of which scaffolding and structure accounted for about 9.5 and the verification loop for about 1.5. Their verifier caught roughly 20% of errors and fixed about 75% of what it caught.
That is a worthwhile margin and we would keep the verifier. But it is a tenth of what structure contributed, at the cost of another model call on every task. It is worth reading as a vendor evaluating its own system, with the limitations the authors list — single scored runs, no same-model scaffold baselines. The directional lesson holds: if you are choosing between adding a judge model and restructuring the task, restructure the task.
7. Batch anything that is not interactive
Asynchronous batch endpoints run at roughly half price across providers. Reconciliation, enrichment, overnight classification, backfills — if no human is waiting, the discount is free money. It is listed last only because it applies to a narrower slice of work.
You cannot manage this without span-level accounting
Every lever above requires knowing which part of which task spent what. Provider dashboards give you a monthly total by API key, which tells you that the bill went up and nothing else. The instrumentation that makes this tractable is not complicated — it is one record per model call, joined to a task identifier.
type AgentSpan = {
taskId: string; // the unit of work you bill against
attempt: number; // which retry this is
stage: string; // checkpoint boundary, for per-stage rates
step: number; // position within the trajectory
model: string;
inputTokens: number;
cachedReadTokens: number; // prove the cache is working
cacheWriteTokens: number;
outputTokens: number;
toolName: string | null;
outcome: "ok" | "tool_error" | "model_error" | "timeout";
latencyMs: number;
};
type TaskOutcome = {
taskId: string;
resolution: "resolved" | "escalated" | "abandoned" | "wrong";
attempts: number;
humanMinutes: number; // the term that dominates the bill
};
With those two tables, the numbers this article is about become ordinary queries: cost per resolved task is the summed span cost joined to outcomes and divided by resolutions. Per-stage success rates fall out of the stage column, which is what tells you where to put the next checkpoint. Cache hit rate falls out of cachedReadTokens over inputTokens, which is how you discover that the cache silently stopped working three deploys ago.
The field most teams omit is humanMinutes, and it is the one that changes decisions. Without it you will optimise the visible 20% of the bill and leave the other 80% alone.
When an agent is the wrong answer
The economics above also tell you when not to build one. Work through these honestly before committing:
| Signal | What it means |
|---|---|
| The steps are known in advance and never vary | You want a workflow with model calls in it, not an agent. Deterministic orchestration is cheaper, testable, and does not drift. |
| A wrong answer is expensive and hard to detect | The escalation term is unbounded. Either add deterministic verification or keep a human in the loop by design. |
| You cannot define "resolved" precisely | You cannot measure the denominator, so you cannot know if it is working. Fix the definition first. |
| The task needs more than ~20 uninterrupted turns | Not impossible, but you are buying the exponential. Decompose or reconsider. |
| Volume is low and each task is high value | The per-task economics may be fine but the engineering never amortises. A good internal tool beats a mediocre agent. |
| The underlying process is broken | Automating it faithfully produces a faster broken process. This is the most common one. |
Mistakes that cost the most
- Estimating from a per-step token count. The quadratic term is not a rounding error; at forty steps it is six times the estimate.
- Benchmarking on clean inputs. Pilot data is curated. Production data has the malformed records, the timed-out API, and the request nobody anticipated — which is where the retry rate lives.
- Treating cost and reliability as separate projects. They are one number with two owners, which is usually why neither owns it.
- Adding a judge model before fixing the structure. An order-of-magnitude difference in return, by the one measurement that separates them.
- Retrying without changing anything. Paying twice for the same deterministic failure.
- Never checking the cache is on. It fails silently and by design, and a single interpolated timestamp disables it.
- No per-task identifier in the logs. Without it, every question in this article is unanswerable.
A sequence that works
- Define "resolved" in a sentence a non-engineer can check. Everything downstream is measured against this.
- Instrument before optimising. Spans joined to task outcomes, including human minutes.
- Measure the baseline on real traffic — cost per attempt, per-stage success, retry rate, escalation rate.
- Turn on caching and verify it with
cache_read_input_tokens. Free, and often the largest single step. - Find the longest uninterrupted trajectory and split it at the first place state can be validated.
- Bound the context by clearing tool results that no later step reads.
- Then, and only then, consider routing and model changes, with the eval set you built from real production failures.
- Re-measure after each change. These levers interact — caching and routing in particular work against each other.
Frequently asked
What is a realistic cost per task for a production agent?
It depends almost entirely on trajectory length and resolution rate, which is why published ranges vary by two orders of magnitude and none of them are about your workload. Calculate it from the formula above using your own step count, prefix size and escalation cost. If you need a sanity check: a well-structured five-to-ten step agent on a mid-tier model lands in cents per attempt, and the human escalation term is usually several times larger than the model spend.
Does a cheaper model reduce cost per resolved task?
Only if it holds the success rate. A model at half the token price with a materially higher retry rate costs more per finished task, because retries re-send the whole context and failures push work into the escalation term. Always compare on cost per resolved task, never on price per million tokens.
How many steps is too many?
There is no universal number, but the compounding is unforgiving: at 95% per-step reliability, twenty steps gets you to 36%. Set a horizon budget per trajectory, measure your own per-step rate, and treat the budget as an architectural constraint rather than a limit to tune upward.
Is prompt caching worth it for short conversations?
Only above the minimum cacheable prefix, which runs from 512 to 4,096 tokens depending on the model. Below it, the request is processed without caching and no error is returned. For agent loops specifically it is almost always worth it, because the stable prefix is re-sent on every step.
Should we build evals before or after shipping?
Build a small set before, from real tasks rather than synthetic ones, and grow it from production failures afterwards. Every incident should become a permanent regression case. The cheapest eval set is the one you harvest from things that already went wrong.
How do we measure the human escalation cost honestly?
Time a sample of escalated tasks end to end — including the context-rebuilding a person has to do because the agent did not hand over cleanly — and load it with the fully burdened hourly cost. It is almost always higher than the estimate, and it is the term with the most leverage.
The short version
An agent's bill is governed by how long it runs without a checkpoint. Token cost grows with roughly the square of the steps; the chance of needing another attempt grows exponentially with them; and the cost of a human cleaning up afterwards dwarfs both. Model price is the lever most teams reach for and close to the least effective one available.
If you are scoping an agent and want the arithmetic run against your actual numbers rather than a published range, send us the workflow and we will work through it with you. If you are earlier than that, taking an AI prototype to production covers the engineering around this one, and our AI and automation work explains how we build these.
Frequently asked questions
What is a realistic cost per task for a production agent?
It depends almost entirely on trajectory length and resolution rate, which is why published ranges vary by two orders of magnitude and none of them are about your workload. Calculate it from the formula above using your own step count, prefix size and escalation cost. If you need a sanity check: a well-structured five-to-ten step agent on a mid-tier model lands in cents per attempt, and the human escalation term is usually several times larger than the model spend.
Does a cheaper model reduce cost per resolved task?
Only if it holds the success rate. A model at half the token price with a materially higher retry rate costs more per finished task, because retries re-send the whole context and failures push work into the escalation term. Always compare on cost per resolved task, never on price per million tokens.
How many steps is too many?
There is no universal number, but the compounding is unforgiving: at 95% per-step reliability, twenty steps gets you to 36%. Set a horizon budget per trajectory, measure your own per-step rate, and treat the budget as an architectural constraint rather than a limit to tune upward.
Is prompt caching worth it for short conversations?
Only above the minimum cacheable prefix, which runs from 512 to 4,096 tokens depending on the model. Below it, the request is processed without caching and no error is returned. For agent loops specifically it is almost always worth it, because the stable prefix is re-sent on every step.
Should we build evals before or after shipping?
Build a small set before, from real tasks rather than synthetic ones, and grow it from production failures afterwards. Every incident should become a permanent regression case. The cheapest eval set is the one you harvest from things that already went wrong.
How do we measure the human escalation cost honestly?
Time a sample of escalated tasks end to end — including the context-rebuilding a person has to do because the agent did not hand over cleanly — and load it with the fully burdened hourly cost. It is almost always higher than the estimate, and it is the term with the most leverage.
Related Posts
Your Context Window Is Not Memory
Every frontier model now advertises a million tokens. Published results show the usable fraction is far smaller, and it moves with the shape of your question. Why bigger windows and higher-recall retrieval both aim the wrong way, and how to measure your own ceiling.
Your Model Is Not a Function
A thousand identical requests at temperature zero produced eighty different answers. The usual explanation about floating point is only half of it — the other half is that your output depends on how many strangers were using the server at the same moment.
Your Agent Has a Permissions Problem, Not a Prompt Problem
Prompt injection is a confused deputy attack, described in 1988 and unfixable at the model layer for the same reason it was unfixable then. The containment architecture already exists, most of it is standardised, and almost nobody ships it.
