Beeija
← Back to Tools

AI Cost Calculators

AI Prompt Caching Savings Calculator

Estimate monthly and yearly savings from reusable prompt prefixes across OpenAI, Claude, Google Gemini, and custom caching prices.

Enter Your Prompt Caching Workload

Separate reusable prompt tokens from the dynamic part of each request.

Cache workload used for this estimate

Cache hits: 85,​000

Cache misses or writes: 15,​000

Cached tokens read: 1,​700,​000,​000

Cache hit rate: 85%

Prompt Caching Cost and Savings

Compare the same request volume with every reusable token billed normally and with the selected cache pricing.

Estimated monthly cost with caching

$525.00

Without caching

$1,​672.50

Monthly saving

$1,​147.50

Saving rate

68.61%

Dynamic input tokens

Always billed at the normal input rate

$37.50

Reusable tokens served from cache

85,000 cache-hit requests

$127.50

Reusable tokens on cache misses

15,000 miss or write requests

$225.00

Output tokens

Output pricing is unchanged by prompt caching

$135.00

Selected pricing: OpenAIGPT-5.4 mini

Estimated yearly saving: $13,​770.00

Cost per request without caching: $0.0167

Cost per request with caching: $0.005250

Approximate break-even hit rate: 0%

Budget status: $475.00 remaining

* Important: Calculated from the values currently shown. Default usage values are examples, so change them to match your expected usage. Built-in OpenAI, Anthropic Claude, and Google Gemini prompt caching rates were checked on June 20, 2026. Final charges may include batch discounts, data residency, priority or fast processing, tool calls, taxes, negotiated discounts, engineering work, and cache misses caused by provider routing or prompt changes.

Long system prompts, tool definitions, examples, documents, and conversation history can be sent repeatedly. Prompt caching can lower the cost of those reusable tokens, but each provider uses a different combination of cache reads, writes, and storage.

Calculating the Cost With and Without Caching

Enter monthly requests, reusable prefix tokens, dynamic input tokens, output tokens, and the expected cache hit rate. The calculator first estimates the cost when every input token is billed at the normal rate.

It then applies the selected provider's cache-read, cache-write, and storage rules. The result shows monthly savings, annual savings, cost per request, savings percentage, and the approximate break-even cache hit rate.

How Provider Caching Models Differ

OpenAI uses automatic prompt caching. Requests with matching prefixes can receive the listed cached-input price, while cache misses use the normal input rate.

Claude charges a higher price when reusable tokens are written to a 5-minute or 1-hour cache, followed by a lower cache-read price when the prefix is reused.

Gemini explicit context caching charges a reduced rate when cached tokens are used and adds storage cost based on cached token volume and time-to-live.

Choosing a Realistic Cache Hit Rate

Use production usage data when available. A stable system prompt shared across many requests can have a high hit rate. Frequently changing instructions, user-specific content, or uneven traffic can lower it.

Cache eligibility and minimum token thresholds vary by provider and model. The calculator assumes the reusable prefix is eligible and structurally identical when a hit is entered.

Practical Decisions This Tool Supports

  • Estimate whether prompt caching is worth implementing.
  • Compare 5-minute and 1-hour Claude cache economics.
  • Model Gemini cache storage and refresh frequency.
  • Estimate savings from long system and tool prompts.
  • Find the approximate cache hit rate needed to break even.
  • Compare current public pricing with a private quote.
  • Plan monthly and yearly LLM cost optimization.

Costs and Limits Not Included

The estimate does not include batch discounts, data residency premiums, priority or fast processing, tool-call charges, taxes, negotiated discounts, latency value, or engineering work.

Cache hits are not guaranteed for every automatic caching system. Prompt structure, token thresholds, retention, routing, request timing, and provider rules can affect real results.

Official Pricing Sources

Built-in prices were checked on June 20, 2026 against the official OpenAI, Anthropic, and Google Gemini API pricing documentation. Mistral's official documentation was also checked for its 10% cached-input rule.

Frequently Asked Questions

What is a prompt cache hit?

A cache hit happens when a request reuses a matching prompt prefix that the provider can serve at its cached-input or cache-read rate.

Why do Claude cache writes cost more than normal input?

Anthropic charges a premium when reusable content is first written to cache. Later cache reads cost a small fraction of the base input rate. The calculator applies the current 5-minute or 1-hour write rate selected.

Why does Gemini add a storage charge?

Gemini explicit context caching bills cached-token reads and also charges for how many tokens are stored and how long each cache object remains active.

Does OpenAI charge a separate cache storage fee?

No separate storage fee is included in OpenAI's automatic prompt caching pricing. Cache misses use the normal input rate and cache hits use the listed cached-input rate.

Does a high cache hit rate always reduce the bill?

Usually, but write premiums, cache storage, short cache lifetimes, changing prompt prefixes, and low reuse can reduce or remove the saving. The calculator shows the break-even hit rate for the entered workload.

Can I model Mistral or another provider?

Yes. Enable custom pricing and enter the provider's current base input, cache-read, cache-write, output, and storage rates. Mistral currently bills cached tokens at 10% of the normal input rate.

Explore Related AI Cost Tools