AI Cost Calculators
Claude API Cost Calculator
Estimate first-party Claude API costs across current models, prompt caching, Batch API, US-only inference, and supported Opus fast mode.
Enter Your Claude API Usage
Use billed token averages from a real request when you have them, then enter the number of requests you expect in one month.
Keep base input, cache writes, and cache reads separate. Output should use the billed output-token count, including thinking tokens when the API reports them.
Rates used per 1 million tokens
Standard · global
Estimated Claude API Cost
First-party Claude API token estimate only; paid server tools, marketplace differences, and other services are separate.
Estimated monthly cost
Per request
$0.005700
Daily avg. (30d)
$7.60
Per year
$2,736.00
Base input cost
36,000,000 tokens
$72.00
Cache write cost
0 tokens
$0.00
Cache read cost
80,000,000 tokens
$16.00
Output cost
14,000,000 tokens
$140.00
Requests: 40,000
Input-related tokens: 116,000,000
Output tokens: 14,000,000
* Important: Built-in first-party Claude API rates checked September 21, 2026. Final charges may include paid server tools, platform-specific pricing, negotiated discounts, taxes, retries, or usage not entered here.
Claude pricing looks simple until one request mixes ordinary input, cached context, internal thinking, a different processing mode, and a residency requirement. Keeping those pieces separate makes the monthly estimate follow the workload you expect to send instead of one headline token price.
Start With the Four Token Buckets Claude Bills Separately
A first-party Claude API request can put usage into four separate billable buckets: base input, cache creation, cache reads, and output. Each bucket has its own rate before the per-request usage is multiplied by monthly request volume.
Processing mode and inference geography are applied after the base model rates. That matters because Anthropic allows pricing modifiers to stack: prompt caching can be combined with Batch processing, and supported models can also use the US-only inference multiplier.
The result is a planning estimate, not a provider invoice. Paid server tools, marketplace billing, taxes, private offers, retries, and services you do not enter here remain outside the total.
The Current Claude Lineup Changes the Cost Comparison
The built-in list follows Anthropic's current model overview: Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5, and Claude Haiku 4.5. Their standard input and output rates are not just scaled versions of one another. Keeping the workload unchanged while switching models shows how much model choice alone changes the monthly total.
There is another reason to re-measure instead of copying an old token count. Anthropic says Claude 4.7 and later models use a newer tokenizer that can produce roughly 30% more tokens for the same text, with the exact change depending on the workload. A prompt measured on an older model is therefore not a dependable token baseline for a newer one.
Fable 5.1 also has a pricing exception worth keeping visible: its cache reads are $0.25 per million tokens, equivalent to 0.025× its base input rate. The other current models listed here use the standard 0.1× cache-read multiplier.
Prompt Caching Has Three Different Billable Paths
A cached prompt is not one generic input bucket. Anthropic prices the uncached portion, the tokens written into a cache, and later cache reads separately. A 5-minute write uses a 1.25× input-price multiplier; a 1-hour write uses 2×.
Cache reads are cheaper, but they only describe tokens that were actually served from an existing cache. Do not enter the same prompt tokens as both base input and cache reads. If your application has a mix of cache hits and misses, use averages from observed API usage or model the two cases separately.
Selecting "No cache write" does not force cache reads to zero. A request can read a cache entry created by an earlier request, so the calculator keeps those two fields independent.
Batch, Fast Mode, and US-Only Inference Solve Different Problems
Batch API is for asynchronous work that can wait. Anthropic currently charges Batch usage at 50% of standard API token rates, and its prompt-caching pricing can stack with that discount.
Fast mode is different. It is a research preview for supported Opus models on the first-party Claude API. It uses premium rates and is designed for higher output-token speed; it is not available together with Batch API. Fast mode only appears here when the selected built-in model supports it. Anthropic currently requires speed: "fast", the fast-mode beta header, and preview access through an account manager or waitlist rather than treating it as a universally available API setting.
US-only inference is another independent modifier. For Claude 4.6 and later models, Anthropic applies 1.1× to input, output, cache writes, and cache reads when inference_geo: "us" is used. Haiku 4.5 is older than that support boundary, so global routing remains the only valid choice here. Inference geography is also separate from workspace data-storage geography; this estimate prices the inference setting, not every residency control.
Use Billed Output Tokens, Not Only the Text You Can See
Claude's billed output can include internal thinking tokens. Anthropic documents output_tokens as the authoritative billed output total, whileoutput_tokens_details.thinking_tokens provides a breakdown for observability.
This matters most when you estimate from screenshots or visible response text. A short visible answer can still have a larger billed output count when the model used more internal reasoning. For production planning, capture actual API usage from representative requests whenever possible.
Build the Estimate From Real API Usage
- Run several requests that resemble the real product flow, including system instructions, tools, retrieved context, and realistic conversation history.
- Record the average base input, cache creation, cache read, and billed output tokens from the API usage fields.
- Enter the expected monthly request volume, then choose the processing mode and inference geography that your actual deployment will use.
- Check a normal month and a busier month rather than relying on one optimistic traffic number.
- If you have negotiated rates or a private commercial agreement, switch to custom standard prices before applying the workload modifiers.
Costs and Boundaries Outside the Token Total
Server-side tools can add their own charges. Web search, for example, is billed separately from model tokens. Tool schemas and tool-result content can also increase normal input usage, so those tokens should be present in the usage figures you enter if they are part of your workflow.
Rate limits, marketplace billing, negotiated discounts, credits, taxes, minimum commitments, and every product-specific charge sit outside this token total. Request-shape checks on the page cover the listed context and output limits, not every validation rule enforced by the Messages API.
Custom prices replace the standard token rates, while the public Batch, fast-mode, and inference-geography multipliers still apply. If a private agreement changes those multiplier rules too, use the effective contracted rates rather than assuming this public-pricing model will match the invoice.
All arithmetic runs in the browser from the numeric usage and price values entered here. There is no prompt or API-key field, and changing a value does not send that value to Anthropic. Opening an official documentation link is a separate browser request to that site.
Official Anthropic References
Pricing checked: September 21, 2026. These links cover the rates and API behaviors that materially change the estimate. Recheck them before a launch or budget approval because model availability and billing rules can change.
- Claude API pricing
Model rates, cache pricing, Batch discounts, residency multipliers, fast mode, and separately billed tools.
- Claude models overview
Current model lineup, model IDs, context windows, and maximum output sizes.
- Prompt caching
Cache-write durations, cache reads, breakpoints, and the behavior behind the caching fields above.
- Batch processing
Asynchronous request behavior and the 50% token discount.
- Data residency
The
inference_geosupport boundary, US-only multiplier, and distinction from workspace geography. - Fast mode
Supported Opus models, gated preview access, the required request setting, and Batch incompatibility.
- Context windows
Input overflow, output interaction, and
model_context_window_exceededbehavior.
Claude Pricing Questions That Change the Math
Why is Claude Fable 5.1 cache-read pricing different?
Anthropic lists Fable 5.1 cache hits and refreshes at 0.025 times its base input rate. The other current models listed here use the usual 0.1 times cache-read multiplier.
Does Batch API also affect prompt-caching charges?
Anthropic states that prompt-caching multipliers stack with Batch API pricing. Batch mode therefore reduces the base input, cache-write, cache-read, and output rates together.
Can Claude Haiku 4.5 use US-only inference?
No. Anthropic documents the first-party inference_geo setting for Claude 4.6 and later models. Haiku 4.5 only supports global routing, so no 1.1 times US-only multiplier applies.
Is fast mode the same as Batch API?
No. Batch is asynchronous processing at discounted token rates. Fast mode is a research-preview option for supported Opus models that uses premium rates for higher output speed. Anthropic does not allow the two modes together.
Why can my billed output tokens exceed the visible reply?
Claude thinking tokens are billed as output tokens even when the full thinking text is not shown. For a real workload, use the API usage output-token count rather than estimating only from visible response text.
Will these numbers match Bedrock or Google Cloud exactly?
Not necessarily. The built-in rates model Anthropic's first-party Claude API. Amazon Bedrock, Google Cloud, Microsoft Foundry, marketplace billing, private offers, and negotiated discounts can use different commercial terms or regional pricing.
