The Cheapest AI API Is Rarely the Cheapest AI Product

The team moved the feature down a model tier because the list price per million tokens was a fifth of what they had been paying. The invoice fell for a month. Then cost per resolved task climbed past where it started: the cheaper model failed the output check more often, every failure was a full-price retry, and the prompt had grown three paragraphs to compensate.

The cheapest AI API by list price is a number on a pricing page. The cheapest AI product is a number on a usage export, and the two disagree more often than not, because list price is one of nine drivers behind cost per request and rarely the largest. The vendor’s pricing page times a token estimate is the comparison most teams run; it prices two of the nine things a request does.

The direct version: the cheapest AI API for a product is the one with the lowest cost per resolved task, which is list price times the token split, adjusted for cache hit rate, retry rate, re-prompt overhead, and tool calls. On the production AI products I have measured, those adjustments added 30 to 80 percent on top of the tokens-times-price estimate, and they moved in opposite directions when a feature changed tier.

LLM pricing comparison: list price versus cost per request

DriverWhat the pricing page showsWhat moves cost per request
Input token pricePer million tokensPrompt length, which grows when a cheaper model needs more instruction
Output token pricePer million, three to five times inputOutput length, which creeps every quarter nobody caps it
Cached input priceRoughly a tenth of uncachedHit rate, which depends on traffic pattern, not the model
Batch discountRoughly half offWhether the workload can wait; most user-facing ones cannot
RetriesNot on the pageFailure rate on the output check, higher on cheaper tiers
Re-promptsNot on the pageQuality-driven second calls, higher on cheaper tiers
Tool callsNot on the pageCalls per run, which rise when a cheaper model plans worse

The first four rows are on the pricing page. The last three are where the model-swap math goes wrong.

The drivers that separate list price from cost per request

Ratios are from vendor pricing pages as of September 2026 (per Anthropic’s pricing page, and the same shape holds at other vendors); the ranges are from production.

List price per million tokens (input rate)

The headline number and the one an AI pricing comparison sorts by. It sets the floor, not the bill. Across features on the same model, cost per request differed ten to one, which means the model price explained less of the spread than the prompt and the product did.

Caveat: two models with the same list price can have different tokenizers, so the same text is a different token count.

Output-to-input price ratio

Output runs three to five times the input rate at the vendors I have priced. On chat assistants, output was 45 to 60 percent of cost per request, so a model that is cheaper on input and the same on output is barely cheaper at all.

Caveat: compare output rates first for chat, input rates first for summarizers.

Cached input pricing (prompt cache reads)

A cache read priced at roughly a tenth of an uncached input token. Caching cut cost per request 20 to 40 percent on chat products with long system prompts and almost nothing on summarizers. A cheaper model without caching can lose to a pricier one with it.

Caveat: a cache write costs more than an uncached read, so content read once loses money in cache.

Batch API discount

Roughly half off for workloads that can wait hours. Applies to classifiers, backfills, and evals; almost never to a user waiting on a response. Two products on the same model with the same list price can differ by half on this alone.

Caveat: a batch job that has to be re-run because it missed its window costs full price twice.

Retry rate (failed output checks)

Second and third calls because the first failed a check: malformed JSON, low confidence, a timeout. Each is a full request at full price. On agents and structured-output features, 10 to 20 percent of cost per request; on plain chat, 5 to 10. The rate climbs when a team moves down a tier.

Caveat: retries are never in the pre-build estimate, so they arrive as unexplained variance at the first invoice.

Quality-driven re-prompts

The second call a product makes because the first answer was not good enough: a “try again” the user clicks, a rewrite step, a verifier that asks for a longer answer. Different from a retry because it is a product decision, and it is the driver that grew most after a tier change.

Caveat: a re-prompt often carries the full history again, so it costs more than the first call.

Tool call count (function calling)

The round trips an agent makes to look something up. One user request fanned out into up to ten model calls; tool calls plus retrieval ran 15 to 30 percent of cost per request on agents. A model that plans worse makes more calls, and each one bills.

Caveat: for agents, cost per request is a distribution. Compare models on the p90 run, not the mean.

Context length (system prompt tax)

The fixed tokens every request carries. On one assistant, the system prompt and tool definitions were more than half of all input tokens for the month. Moving to a cheaper model usually means adding instructions, which is a price increase on every call.

Caveat: the cheapest lever on the list and the one skipped because it looks like a chore.

Pay-per-token versus provisioned throughput

The cheapest API on a per-token page can be the most expensive on a provisioned pool. At low utilization, a request on a provisioned pool ran two to three times the pay-per-token equivalent; at high utilization, below it. The comparison depends on how the capacity is bought.

Caveat: a request on a provisioned pool has no marginal cost until the pool fills, so it looks free until it looks like a resize.

Eval overhead (LLM-as-judge)

The second model grading the first. Sampled at 5 to 10 percent of traffic it stayed at a few percent of cost per request; run on every request it nearly doubled it. A cheaper primary model that needs a pricier judge on every call has not saved anything.

Caveat: eval spend lands on the platform team’s bill, not the product’s, unless someone allocates it.

Cost per resolved task (cost per successful outcome)

The number the comparison should end on: total spend divided by tasks the product completed, retries and re-prompts included. On structured-output features, the cheaper model cost more per resolved task after its retry rate rose, which is the whole argument in one metric.

Caveat: “resolved” needs a definition the eval can check, and the eval has its own cost.

The model-swap math: what changes when a feature moves down a tier

Four things move at once, and only the first is on the pricing page. List price falls. Retry rate rises, because the cheaper model fails the output check more often. Prompt length rises, because the team adds instructions to close the quality gap. Re-prompts rise, because users and verifiers ask for a second answer more often. On production products the first effect was often smaller than the other three combined.

When the cheaper tier actually wins

It wins on classifiers, extraction, and routing: short prompts, short outputs, a check the model passes at nearly the same rate as the pricier one. It wins on any workload that can go through the batch discount. It loses on agents, on long-context chat, and on anything where the output check fails more than a few points more often, because each failure is a full-price call.

Price per token versus price per task: which to compare on

Compare on cost per resolved task at the product’s real token split, cache hit rate, and retry rate, at three usage tiers. Price per token is the input, not the result. An AI API cost estimation that stops at price per token estimates the cheapest invoice, not the cheapest product.

Pre-build versus in production: where the comparison comes from

Before the build, the comparison runs in the LLM cost calculator with assumed splits and rates, good for narrowing to two candidates. After launch it comes from the usage export, per feature, with retry and re-prompt rates measured. Run both models on the same traffic sample for a week before the swap, and compare cost per resolved task, not the invoice.

The mistake: swapping on list price and reporting the invoice

The mistake is moving a feature down a tier because the list price is lower, then reporting the lower invoice as the saving. The invoice is lower for as long as it takes the retry rate and the prompt length to catch up. The symptom: a cost line that drops, then climbs back over two quarters with traffic flat, and a quality complaint in between. On a Fortune-500 financial data platform with eight figures of annual AI and cloud spend, the feature-level breakdown made this visible; the invoice never would have.

The rule

Compare AI APIs on cost per resolved task at your token split, cache hit rate, and retry rate, never on price per million tokens. The pricing page is the first input to that number and the only one most comparisons use.

Per-token prices, cache and batch discounts, and rate limits change monthly. The ratios above are dated September 2026 and should be checked against each vendor’s live pricing page before they go into a decision.

FAQ

What is the cheapest AI API?

By list price, whichever small model tops the comparison this month, which changes often enough that no article should name one. By cost per request, the one with the lowest price times your token split, adjusted for cache hit rate, retries, re-prompts, and tool calls; on production products those adjustments added 30 to 80 percent and reordered the list.

How do you compare LLM pricing across vendors?

Build cost per request for your feature at its measured token split: input tokens times the input rate, output tokens times the output rate (three to five times input), cached reads at roughly a tenth, plus retries and tool calls at full price. Run it at three usage tiers. Then divide by resolved tasks, because the cheaper model’s retry rate is where the comparison usually flips.

Is a cheaper model always cheaper to run?

No. A cheaper model that fails the output check more often costs more per resolved task, because each failure is a full-price retry and the prompt tends to grow to compensate. On structured-output features in production, the cheaper tier cost more per resolved task after its retry rate rose. Cheaper models win on classifiers, extraction, and batch workloads.

What to do next

Take one feature, run a week of its traffic through both candidate models, and compare cost per resolved task instead of the invoice. The nine drivers behind that number are in what one AI request actually costs, and the monthly version is in how much AI costs per month by product shape. When the comparison needs to become a pricing decision finance signs, that is FinOps consulting for AI products; the rest of the FinOps for AI practice is on the home page.


I put a number on what an AI product costs per request, per active user, and per subscriber, early enough to change the model, the architecture, or the price. If you’re shipping something with a model behind it and nobody can tell you what it costs, email me.

Have a number nobody can explain?

Send a note and I will get back to you.