Anthropic API costs: measure the cost of a successful AI task
Calculate Claude API spend, account for cache writes and retries, and turn cost per successful task into a realistic budget for your app's AI feature.
By AppCradle · Published

The useful cost of an AI feature is the amount you spend to deliver an acceptable result. A cheap model response can still be expensive if the app needs another attempt, a fallback model or a person to fix it. Start with a billable request, then work toward the successful task and the customer account that pays for it.
For an app that summarizes documents, for example, define success as a delivered summary that passes your factual and formatting checks. Receiving an HTTP response is a technical outcome; it does not establish that the customer received a usable summary.
This article uses a dated Claude API price snapshot and explicitly hypothetical workloads. The examples are calculations, not AppCradle customer results or a forecast of your bill. All amounts are in US dollars.
Build a rate card for the route you actually use
As checked on October 6, 2026, Anthropic lists the following Claude Sonnet 5.5 rates for the direct API. This snapshot uses standard processing and global inference, without negotiated discounts, taxes or separately billed tools. Recheck the official pricing page before applying it to a budget.
| Billing category | USD per million tokens |
|---|---|
| Uncached input | $2.00 |
| Output | $10.00 |
| Five-minute cache write | $2.50 |
| One-hour cache write | $4.00 |
| Cache read | $0.20 |
Store the effective date, model identifier, provider route and processing configuration beside your rates. A model name alone is not a complete billing specification. If your application uses a cloud partner, a regional route or another service tier, use that route's applicable price and billing records.
Use token counting to estimate supported inputs before a request, then reconcile against observed usage. Anthropic describes its token count as an estimate; it is not a prediction of the eventual response length. Recount representative inputs when changing models instead of assuming the same text has the same token count everywhere.
Keep cache categories separate
Anthropic's cache usage documentation distinguishes ordinary input, cache creation and cache reads. In a cached request, input_tokens excludes tokens counted in the other two categories. Do not charge the combined input total at the ordinary rate and then add the cache charges again.
For one request, calculate:
Token cost = (
uncached input tokens × input rate
+ 5-minute cache-write tokens × 5-minute write rate
+ 1-hour cache-write tokens × 1-hour write rate
+ cache-read tokens × read rate
+ output tokens × output rate
) / 1,000,000
The rates in this formula are per million tokens. When both cache durations are used, take the duration-specific creation counts from the usage breakdown. They partition the cache-creation total; they are not extra tokens to add on top of it.
Here is a hypothetical sequence using the price snapshot above. Each request has 8,000 input tokens and 600 output tokens. Of the input, 6,000 tokens form an unchanged, cache-eligible prefix. Assume the first request writes a five-minute entry and the next nine requests all hit it before expiry.
| Request pattern | Calculation | Token cost |
|---|---|---|
| One uncached request | (8,000 × 2 + 600 × 10) / 1,000,000 | $0.0220 |
| First request, writing the prefix | (2,000 × 2 + 6,000 × 2.50 + 600 × 10) / 1,000,000 | $0.0250 |
| A later cache hit | (2,000 × 2 + 6,000 × 0.20 + 600 × 10) / 1,000,000 | $0.0112 |
| Ten requests without caching | 10 × 0.0220 | $0.2200 |
| One write and nine hits | 0.0250 + 9 × 0.0112 | $0.1258 |
Those savings depend on the assumed hits. Do not put them into a forecast merely because caching is enabled. Measure actual writes and reads, check the model's minimum cacheable length, and test the real spacing between requests. A frequently changing prefix or a long gap can invalidate the planning assumption. See Anthropic's cache limitations and lifetime guidance.
Keep this comparison at the token-cost level. It does not prove that the whole feature became cheaper by the same proportion.
Count all attempts against the original task
Give each user task an internal identifier and associate every model call, retry and fallback with it. Keep a separate identifier for each attempt so that repeated log delivery does not create duplicate costs. Protect those operational identifiers; they do not belong in public analytics events.
A minimal internal record should answer four questions:
- What ran: model, prompt version, provider route and processing configuration.
- What was consumed: usage categories and separately attributable tool or infrastructure costs.
- What happened: completed, rejected by validation, failed, abandoned or awaiting review.
- Whether it helped: the task's acceptance result, recorded once under a documented rule.
A validation failure can follow a perfectly valid, billable model response. A transport failure may leave usage unknown. Keep unknown costs visible until reconciliation; neither assume every error was billed nor replace missing usage with zero. Anthropic's error documentation describes request identifiers that can help investigate individual calls.
Bound retries and distinguish temporary failures from outputs that need a different strategy. Sending the same request indefinitely can increase cost without improving the result. For tools that change external state, protect the action against duplication as well as protecting the cost record.
Calculate cost per successful task
Choose a cohort with enough time for its tasks to finish, then count outcomes using the same acceptance rule. Include attributable spending on unsuccessful tasks in the numerator:
Cost per successful task =
total attributable cost for the cohort
/ number of accepted tasks in that cohort
If no tasks pass, report zero successful tasks and the total cost; the ratio is undefined. If work is still pending, mark the cohort provisional instead of mixing its unfinished spending with a final success count.
Consider a second hypothetical example with caching disabled. There are 1,000 started tasks and 1,200 billable model calls, including retries. Assume each call averages the 8,000 input and 600 output tokens used above. Only 900 tasks meet the acceptance rule.
| Cohort cost | Assumed amount |
|---|---|
Model calls: 1,200 × $0.022 | $26.40 |
| External tool charges | $4.00 |
| Attributable infrastructure | $6.00 |
| Human review allocation | $9.00 |
| Total | $45.40 |
Model spend per successful task is $26.40 / 900, or about $0.0293. With the listed supporting costs, it becomes $45.40 / 900, or about $0.0504. Quoting only the $0.022 average call price would hide both repeat work and the surrounding service cost.
Use the actual billable workload, not a count of clicks. Treat the external-tool and human-review amounts here as explicit assumptions, not provider prices. Document your allocation method and avoid counting any tool charge both in a provider cost report and a separate ledger.
Translate the result into a plan budget
Start with comparable net proceeds for a paying account and reserve room for the rest of the service. Do not compare a tax-inclusive customer price with a cost figure while ignoring store deductions or refunds. The sales versus proceeds guide explains the revenue-side distinction.
For illustration, assume one account produces $10 of monthly net proceeds and incurs $3 of other costs, excluding everything already allocated above. At the example's average cost, 100 successful tasks would consume about $5.04, leaving about $1.96 before other unallocated expenses. This is a scenario, not a safe universal allowance.
The average is only a starting point. Examine long documents, repeated retries, output-heavy tasks and your most active accounts separately. A task allowance is meaningful only if the task definition also bounds expensive inputs, outputs and tool activity. Decide how the product explains an exceeded allowance before enforcing it.
If the economics fail, test a concrete change: shorten unnecessary context, reduce redundant attempts, limit output to what the feature needs, or route a defined subset of tasks to another model. Compare quality and completion alongside cost. Our monetization experiment guide helps frame that decision without attributing every revenue change to the experiment.
Anthropic's Message Batches API offers discounted asynchronous processing. Consider it for workloads that can wait; check completion, expiration and result-handling requirements before changing an interactive feature to a background job.
Reconcile before treating the estimate as settled
Keep two views: an operational estimate for individual tasks and the provider's reported usage and costs. Compare the same dates, workspace, model and billing scope. Investigate differences such as delayed records, duplicate ingestion, unallocated requests or a stale rate card.
Anthropic provides a Usage and Cost Admin API for eligible organizations with appropriate admin credentials; it is not available to individual accounts. Its Cost API reports daily amounts as decimal strings in US cents. Convert units explicitly before comparing them with a dollar-denominated spreadsheet. Keep privileged credentials on the server.
Use aggregates for a cost review. Raw prompts and model responses are unnecessary for arithmetic and may contain sensitive customer data. Keep any quality-review samples under a separate access and retention policy.
AppCradle's store and subscription reports can inform the proceeds side of this review. The Anthropic usage ledger and task outcomes described here must come from your own systems; this article does not describe an AppCradle Anthropic integration. Compare aligned dates and currencies, using the revenue reconciliation guide where needed.
Bring the rate-card date, total cost, accepted-task count, expensive workload segments and unresolved billing differences to your weekly metrics review. Those are the inputs for deciding whether to change the implementation, the allowance or the price of the feature.
See the bigger picture.
Bring your store reports together and spend more time understanding what they mean.
Try for free