What an AI Agent Costs to Run: Cost Per Resolved Task, Not Per Request

The agent shipped, and Finance asked what a task costs. Engineering’s back-of-envelope priced it the way the team had priced the chat feature: one prompt in, one answer out, times the model’s per-token rate. The usage export disagreed. The loop that resolves one ticket ran eleven steps on the median case and forty on the worst one, and every step was a full call at full price.

That gap is the whole problem with ai agent cost. A request has one call and one price. An agent plans, calls a tool, reads the result, and decides again, and its context grows by one more turn each time, so the twentieth call pays for the first nineteen. Priced like a single request, the way most launch estimates get built, the number is wrong before the first re-plan.

An AI agent’s cost is not one request; it is a loop, and the loop decides the price. Cost per resolved task is steps per task times the tokens each step carries, plus tool-call payloads, retries, dead ends, and verification passes. Per Anthropic’s engineering post on its production Research system, an agent runs at roughly four times a single chat interaction’s token count.

Twelve of the drivers below only exist once a request becomes a loop. None of them show up on the invoice, and none of them are in a launch estimate built from the model’s list price alone.

AI agent cost: the twelve drivers

Same profile for each: what it is, what moves it, the range it landed in, the caveat. The ranges come from an agentic workflow inside a Fortune-500 financial data company’s production AI portfolio, abstracted to shape.

Steps per task (the loop count)

The number of plan-act-observe cycles one task takes to resolve. The single biggest lever on cost, because every other driver on this list multiplies by it. On a support-resolution agent, the median task ran 8 to 12 steps; the p90 ran past 35.

Caveat: step count is a distribution, not a constant, so one “steps per task” number in a launch estimate hides the tail that drives the bill.

Context growth per step (the compounding prefix)

Each step appends its output, tool result, and reasoning to what the model reads next. A ten-step task doesn’t read ten prompts; it reads one that grew by nine turns of history. Input tokens on a long-running task ran 3 to 6 times the first step’s alone.

Caveat: a per-request cost model can’t see this driver, because per-request pricing treats every step as identical.

Tool-call tokens (function-calling overhead)

The schema, arguments, and return payload for every tool invocation, priced as ordinary input and output tokens. A single web-search or database call can return more tokens than the user’s original question. On the workflow measured, tool payloads ran 20 to 35 percent of task cost.

Caveat: a verbose tool response, a full payload instead of a filtered one, is a pricing decision disguised as an engineering one.

Retries and dead ends

Steps that fail a check, hit a malformed tool call, or explore a path the agent abandons. Each is a full-price call with no forward progress. On the workflow measured, dead ends added 10 to 25 percent on top of the steps that resolved the task.

Caveat: dead ends are invisible in a step-count average; they only show up when you measure spend per resolved task instead of per attempted one.

Verification and self-check passes

A second pass, often a second model call, that checks the agent’s own output before it ships: does this answer the question, does this tool result look right. Sampled lightly, verification adds a few percent; run after every step, it can nearly double task cost.

Caveat: skipping verification to save cost is how a cheaper agent gets more expensive once wrong answers turn into support tickets.

Sub-agent orchestration (fan-out)

A lead agent that spins up multiple subagents to work a task in parallel, each with its own context window, tools, and token budget. Per Anthropic’s engineering post, direct-comparison tasks used 2 to 4 subagents at 10 to 15 tool calls each; complex research used more than 10.

Caveat: fan-out buys speed and thoroughness, not savings; every subagent is its own token budget, not a shared one.

Planning and re-planning overhead

The step where the agent decides what to do next, separate from the steps that do it. A task that changes direction mid-run, a tool returns something unexpected, a subagent’s findings contradict the plan, pays for a re-plan on top of the work already done.

Caveat: this cost rises with task ambiguity, which is exactly the case a launch estimate predicts worst.

Tool schema tokens (the fixed-definitions tax)

The tool definitions and descriptions sent with every call, whether or not that call uses the tool. An agent with 15 tools defined pays for describing all 15 on every step.

Caveat: the cheapest lever on this list, and the one most skipped, because trimming an unused tool looks like housekeeping, not a pricing decision.

Parallel tool calls (concurrent branches)

Multiple tool calls issued in the same step instead of sequentially, which cuts latency but not token cost; each call is still priced on its own. Four tools fired in parallel cost the same as four fired in sequence.

Caveat: teams that added parallel calls to hit a speed target were often surprised the invoice didn’t move, because parallelism is a latency fix, not a cost fix.

Long-horizon runs (session length)

Tasks that stay open across many turns of user interaction, not just many internal steps: a multi-day research project, an ongoing monitoring agent. Cost accrues continuously, so “cost per task” needs a time boundary or it never closes.

Caveat: a long-horizon agent without a session boundary produces a number that only ever goes up, which is useless for a monthly run rate finance can plan against.

Cache hit rate on a growing context

The share of the agent’s ballooning prefix served from cache rather than priced at the full input rate. Per Anthropic’s pricing page as of September 2026, a cached read runs at roughly a tenth of the uncached rate, and the discount matters more here, because the prefix an agent re-reads is larger on every step.

Caveat: the cache breaks the moment something before the breakpoint changes, and an agent’s context changes by design every step, so its hit rate needs its own measurement, not the chat product’s.

Stop-rule failures (runaway loops)

A task that never resolves cleanly: the agent loops between two tool calls, or keeps re-planning without converging. Without a hard step cap, one runaway task can cost as much as several hundred resolved ones.

Caveat: this is the tail risk a launch estimate never models, because a model built on the median task has no term for one that doesn’t end.

AI agent cost by task complexity

Anthropic’s engineering post on its production Research system names three task-complexity buckets and the agent and tool-call counts each one used. It doesn’t publish a cost multiplier per bucket; separately, across its production traffic as a whole, it reports agents running at roughly four times the tokens of a single chat interaction and multi-agent systems at roughly fifteen times. Read the two together, not as one table Anthropic published, but as the shape a launch estimate should expect:

Task complexityAgentsTool calls
Simple fact-finding1 agent3 to 10
Direct comparison2 to 4 subagents10 to 15 each
Complex, multi-step research10+ subagentsclearly divided, no fixed count

Token usage by itself explained roughly 80 percent of the performance variance on Anthropic’s browsing benchmark, ahead of tool-call count and model choice combined. Budget by task complexity first, then by model: the agent-and-tool-call shape moves cost more than which model runs it, the same point the cheapest AI API is rarely the cheapest AI product makes about a single request.

From cost per step to cost per resolved task

Cost per request has one hop to cost per user. Cost per step has three.

Cost per step. One plan-act-observe cycle’s tokens, tool payload, and cache-adjusted price. The atomic unit, useful for finding which step type in a workflow runs expensive.

Cost per resolved task. Every step’s cost, across every attempt including dead ends and retries, divided by tasks that actually completed. This is the number to report; a task that took forty steps and succeeded is not the same line item as one that took forty and gave up.

Cost per completed workflow. For a multi-task process, an agent that researches, then drafts, then verifies, the sum of every task’s cost across the chain. Workflows are what a business case is built on, not tasks in isolation.

Steps-to-completion ratio. Steps attempted divided by tasks resolved. A ratio climbing while step count stays flat means more tasks are hitting retries or dead ends, the earliest signal that cost per task is about to rise.

Bounding agent cost: step caps and stop rules

An agent without a bound on it doesn’t have a cost; it has a distribution with no upper tail. Four rules keep that tail off the invoice.

Set a hard step cap per task, sized to the measured p90, not a guess. Anthropic’s own agents cap simple fact-finding around 3 to 10 tool calls for the same reason: bound the loop before it needs bounding.

Cap verification, don’t skip it. Sample self-checks at a fixed rate, 5 to 15 percent of tasks, rather than running one on every step or none at all; both ends of that range cost more than the middle, just in different currencies, tokens versus wrong answers.

Budget subagents like headcount, not a free parallel call. Each one carries its own context window and token cost, so fan-out width is a design-time decision, not something the agent decides unbounded at runtime.

Give every long-horizon run a session boundary. A cost number with no time limit isn’t a metric, it’s a running total. Close the session, price it, start a new one.

The mistake: pricing an agent like a single request

Take the chat feature’s cost-per-request number, apply it to the agent, and call it done. It’s low by the token multiple a loop adds, because a single-request estimate has no term for steps, tool payloads, retries, or the compounding context that makes step twenty cost more than step one.

The symptom is a business case that looked fine at launch and a margin that quietly went negative by month two, once real tasks hit the tail of a step-count distribution the estimate never modeled. The fix is upstream of launch: the AI business case a CFO will sign needs cost per resolved task in it, not cost per request with an asterisk.

The rule

Price an agent from its steps, cap the loop before you ship it, and report cost per resolved task, never per request. The cap clause matters as much as the pricing clause: an unbounded loop doesn’t just cost more, it makes the number itself meaningless, because a distribution with no upper limit has no single price to report.

What to do next

Take one agentic feature, pull a week of runs, and build cost per resolved task from steps, tool calls, and retries instead of the model’s list price. The LLM cost calculator is a starting point for the token-level inputs; turning that into a step-capped forecast Finance can run unassisted is the work on FinOps consulting for AI products. The rest of the FinOps for AI practice is on the home page.

FAQ

How much does an AI agent cost to run?

There’s no single number; it depends on steps per task, not the model’s per-token rate alone. A simple, single-agent task with 3 to 10 tool calls costs a few times a chat request; a multi-agent task with 10 or more subagents can run fifteen times that, per Anthropic’s own production data.

What’s the difference between AI agent cost and LLM cost per request?

A request is one call, priced from its own drivers: input, output, and cached tokens, plus retries. An agent is a loop of many requests, where each step’s context includes everything before it, so cost compounds with steps instead of staying flat per call.

How many tool calls does an AI agent need?

It scales with task complexity. Anthropic’s production data puts simple fact-finding at 1 agent with 3 to 10 tool calls, direct comparisons at 2 to 4 subagents with 10 to 15 calls each, and complex research above 10 subagents with no fixed count, so the tool-call budget belongs in the estimate before the model choice does.

How do you bound the cost of an AI agent?

A hard step cap sized to the measured p90, a capped (not skipped) verification rate, a fixed subagent budget set at design time, and a session boundary on any long-running task. Without those four, an agent’s cost isn’t a number, it’s an open-ended distribution.


I put a number on what an AI product costs per request, per active user, and per subscriber, early enough to change the model, the architecture, or the price. If you’re shipping something with a model behind it and nobody can tell you what it costs, email me.

Have a number nobody can explain?

Send a note and I will get back to you.