Beeija
Home/Tools
← Back to Tools

AI Cost Calculators

DeepSeek API Cost Calculator

Estimate current DeepSeek API spend by combining token usage, cache behavior, and the provider's time-of-day pricing windows.

Estimate your DeepSeek workload

Choose the model, enter one representative request, then describe how much input is served from cache and how much traffic runs outside DeepSeek's weekday peak windows.

Weekday UTC pricing clockWeekends: off-peak all day
000104061024

Yellow marks DeepSeek's weekday peak windows; green marks off-peak.

Current token rates

DeepSeek V4.1 Flash · API deepseek-flash · USD per 1 million tokens.

Token pathPeakOff-peak
Cache-hit input$0.006000$0.003000
Cache-miss input$0.30$0.15
Output$1.20$0.60

Monthly DeepSeek estimate

Cache-hit input, cache-miss input, and output are priced separately, then blended using the off-peak share you entered.

Estimated monthly cost

$68.616

Per request

$0.000858

Daily average

$2.2872

12 months at this mix

$823.392

Cache-hit input$0.216
Cache-miss input$25.20
Output$43.20

Input tokens: 120,000,000 · cache hit 36,000,000 · cache miss 84,000,000

Output tokens: 36,000,000 · off-peak share 0%

Effective rates / 1M tokens: hit $0.006000, miss $0.30, output $1.20

API model: deepseek-flash · published account concurrency limit 2,500

* Important: Calculated from the values currently shown. Default usage values are examples, so change them to match your expected usage. Built-in DeepSeek rates were checked on September 22, 2026. Final charges may include taxes, credits, retries, tool-side services, account-specific terms, and usage outside the token rates entered here.

Off-peak pricing matters only when the workload can actually move

DeepSeek's weekday peak windows are 01:00–04:00 and 06:00–10:00 UTC. Traffic outside those windows, plus all weekend traffic, uses the lower off-peak rates shown in the calculator. The pricing clock beside the off-peak field is there to make that schedule visible while you are choosing the workload share.

Delayed evaluation, indexing, summarization, and other background jobs may be movable into cheaper hours. User-facing requests usually are not. If the workloads also have very different prompt or output sizes, estimate them separately rather than blending unlike traffic into one percentage.

Cache savings should come from token usage, not request counts

DeepSeek's context cache is enabled by default, but a repeated request does not mean the whole prompt is billed as a cache hit. The matching prefix must already be available to the cache, and the cache operates on a best-effort basis.

For live traffic, the useful measurements are the token counters returned by the API. A request-level hit rate can distort the budget when some prompts are much larger than others, so the calculator's cache percentage represents the share of input tokens billed at the cache-hit rate.

prompt_cache_hit_tokens
Input tokens billed at the cache-hit price.
prompt_cache_miss_tokens
Input tokens that missed the cache and use the higher input rate.
Cache-hit share in the calculator
A planning shortcut until measured token totals are available; replace it with observed usage when you have it.

V4.1 Flash changes the name, while billing still depends on usage

The current Flash model is deepseek-flash, which DeepSeek identifies as V4.1 Flash. The older deepseek-v4-flash alias remains accepted for compatibility, but routes to V4.1 Flash and uses the current Flash pricing. deepseek-v4-pro remains separately priced.

Both models are documented with a 1 million-token context window and maximum output of 384,000 tokens. The calculator rejects an average request already beyond those published boundaries instead of returning a misleading cost for an impossible request shape.

Thinking tokens belong in the output budget

Thinking is enabled by default in the current model family unless it is disabled. Reasoning can be exposed separately for observability while still contributing to billed output usage.

For a production estimate, use the API's returned output-token usage rather than estimating cost from the visible answer alone. In Responses API usage, reasoning can be broken out through output_tokens_details.reasoning_tokens, while total output usage remains the billing input that matters here.

Before trusting the monthly total in production

Concurrency

DeepSeek currently documents account-level concurrency limits of 2,500 for deepseek-flash and 500 for deepseek-v4-pro. Exceeding the applicable limit can return HTTP 429 even when the token budget is acceptable.

User isolation

The optional user_id mechanism affects cache and scheduling isolation. DeepSeek says not to place private user information in that identifier, so use an internal non-sensitive identifier when the feature is needed.

Outside the estimate

Taxes, granted balance, credits, retries, surrounding services, hosting, network costs, account-specific terms, and future price changes are outside the token estimate.

The arithmetic runs locally in your browser. There is no prompt or API key field, and changing a workload value does not send it to DeepSeek.

DeepSeek sources used for this page

Pricing checked: September 22, 2026.Current rates, model names, cache behavior, token accounting, thinking behavior, and request limits were checked against DeepSeek's own documentation: models and pricing, context caching, token usage, thinking mode, rate limits and isolation. Recheck the provider pages before a launch or budget approval because pricing and limits can change.

Explore related AI cost tools