Beeija
Home/Tools
← Back to Tools

AI Cost Calculators

OpenAI API Cost Calculator

Estimate text-token costs for current OpenAI models using uncached input, cached input, cache writes, output tokens, Standard or Batch processing, and custom rates.

Enter Your OpenAI API Usage

Enter average token usage for one request, then add the number of requests you expect in a month.

Cached input means tokens read from an existing prompt cache. Cache writes are tokens newly written to cache on GPT-5.6 and later models.

Rates used per 1 million tokens

Standard

Uncached input:$10.00

Cached input:$1.00

Cache write:$12.50

Output:$50.00

Estimated OpenAI API Cost

Text-token estimate only; tool calls and other paid services are separate.

Estimated monthly cost

$1,160.00

Per request

$0.0232

Daily avg. (30d)

$38.6667

Per year

$13,920.00

Uncached input

40,000,000 tokens

$400.00

Cached input

10,000,000 tokens

$10.00

Cache writes

0 tokens

$0.00

Output

15,000,000 tokens

$750.00

Requests: 50,000

Input tokens per request: 1,000

Total input tokens: 50,000,000

Total output tokens: 15,000,000

* Important: Built-in rates checked September 16, 2026. Final OpenAI charges can also include tool calls, media, storage, regional processing, taxes, discounts, or usage not entered here.

OpenAI API billing is not just one token rate. The model, uncached input, cached reads, cache writes, output length, processing mode, and long-context rules can all change the result. This calculator keeps those pieces separate so the estimate is easier to inspect.

It calculates locally in your browser and does not need an API key. ChatGPT subscriptions are separate from OpenAI API billing.

What This Calculator Includes

The built-in model list focuses on OpenAI's current flagship API models: GPT-6 Astra and the GPT-5.6 Sol, Terra, and Luna family. For each request, enter uncached input, cached input, cache-write tokens, and output tokens separately.

The calculator multiplies those token totals by the selected per-million-token rates, then shows estimated cost per request, a 30-day daily average, monthly cost, and yearly cost. Custom pricing lets you replace the built-in base rates without changing the workload.

Prompt Caching: Reads and Writes Are Different

A cached input token is a token read from a reusable prompt prefix and billed at the lower cached-input rate. A cache write is different: it is a token being written into the prompt cache for possible reuse later.

OpenAI currently charges cache writes on GPT-5.6 models and later model families at 1.25× the uncached input rate. The API reports cache reads in cached_tokens and writes in cache_write_tokens. Enter the average values you actually observe when possible instead of assuming every repeated prompt will become a cache hit.

Prompt caching is automatically available for eligible prompts, but a shared prefix must match for a cache read to occur. Changing content inside the reusable prefix can turn an expected hit into another write or an uncached request.

Batch and Long-Context Pricing

Batch API is intended for asynchronous work that does not need an immediate response. OpenAI documents Batch as 50% lower cost than synchronous APIs, with batches completing within a 24-hour window. Select Batch only when the workload can actually use that processing path.

The current GPT-6 Astra and GPT-5.6 family also have a long-context pricing rule. When a request exceeds 272,000 input tokens, this calculator automatically applies 2× to input and cache rates and 1.5× to output rates for the full request.

The long-context check uses uncached input, cached input, and cache-write tokens together because all are part of the request's input-token volume.

A Better Way to Build the Workload Estimate

Start from measured or realistic average usage rather than a smallest-case prompt. Include system instructions, user messages, conversation history, retrieved context, tool descriptions, and other text that becomes model input.

Keep cache reads and cache writes separate. A mature workload with a stable reusable prefix may have a very different cost profile from a new or frequently changing prompt that keeps writing fresh cache entries.

For output, use the answer length you expect in production. A classification task, coding agent, support assistant, and long-form report generator can have very different output token usage even at the same request count.

Worked Example

Suppose an application makes 50,000 requests in a month. An average request has 800 uncached input tokens, 200 cached input tokens, no cache write, and 300 output tokens. Enter those values, choose a model, and the calculator separates each token category before adding the monthly total.

To compare models fairly, leave the workload unchanged and change only the model. If you are comparing Standard with Batch, keep the token assumptions the same so the processing mode is the only changing variable.

What the Estimate Does Not Include

This is a text-token estimate. It does not add separate charges for web search, file search, image generation, audio, video, containers, storage, code execution, or other paid tools and services that may be used alongside a model.

It also does not add regional-processing or data-residency uplifts, taxes, account-specific discounts, Scale Tier or other contracted capacity, credits, retries that are not included in your request count, or provider changes made after the checked date.

GPT-5.6 Sol's current listed API rate is promotional according to OpenAI and is stated as available at least through November 21, 2026. Re-check that rate before using the result for a budget that extends beyond the promotion.

Pricing Sources and Checked Date

Built-in rates and billing rules were checked against official OpenAI documentation on September 16, 2026. OpenAI can change models, rates, caching rules, service tiers, and other billing behavior after that date.

Review the OpenAI model catalog ↗, prompt caching guide ↗, and Batch API guide ↗ before making a final budget or purchase decision.

Explore Related AI Cost Tools