AI Cost Calculators
DeepSeek API Cost Calculator
Estimate current DeepSeek API spend by combining token usage, cache behavior, and the provider's time-of-day pricing windows.
Estimate your DeepSeek workload
Choose the model, enter one representative request, then describe how much input is served from cache and how much traffic runs outside DeepSeek's weekday peak windows.
Yellow marks DeepSeek's weekday peak windows; green marks off-peak.
Current token rates
DeepSeek V4.1 Flash · API deepseek-flash · USD per 1 million tokens.
Monthly DeepSeek estimate
Cache-hit input, cache-miss input, and output are priced separately, then blended using the off-peak share you entered.
Estimated monthly cost
Per request
$0.000858
Daily average
$2.2872
12 months at this mix
$823.392
Input tokens: 120,000,000 · cache hit 36,000,000 · cache miss 84,000,000
Output tokens: 36,000,000 · off-peak share 0%
Effective rates / 1M tokens: hit $0.006000, miss $0.30, output $1.20
API model: deepseek-flash · published account concurrency limit 2,500
* Important: Calculated from the values currently shown. Default usage values are examples, so change them to match your expected usage. Built-in DeepSeek rates were checked on September 22, 2026. Final charges may include taxes, credits, retries, tool-side services, account-specific terms, and usage outside the token rates entered here.
Off-peak pricing matters only when the workload can actually move
DeepSeek's weekday peak windows are 01:00–04:00 and 06:00–10:00 UTC. Traffic outside those windows, plus all weekend traffic, uses the lower off-peak rates shown in the calculator. The pricing clock beside the off-peak field is there to make that schedule visible while you are choosing the workload share.
Delayed evaluation, indexing, summarization, and other background jobs may be movable into cheaper hours. User-facing requests usually are not. If the workloads also have very different prompt or output sizes, estimate them separately rather than blending unlike traffic into one percentage.
Cache savings should come from token usage, not request counts
DeepSeek's context cache is enabled by default, but a repeated request does not mean the whole prompt is billed as a cache hit. The matching prefix must already be available to the cache, and the cache operates on a best-effort basis.
For live traffic, the useful measurements are the token counters returned by the API. A request-level hit rate can distort the budget when some prompts are much larger than others, so the calculator's cache percentage represents the share of input tokens billed at the cache-hit rate.
prompt_cache_hit_tokens- Input tokens billed at the cache-hit price.
prompt_cache_miss_tokens- Input tokens that missed the cache and use the higher input rate.
- Cache-hit share in the calculator
- A planning shortcut until measured token totals are available; replace it with observed usage when you have it.
V4.1 Flash changes the name, while billing still depends on usage
The current Flash model is deepseek-flash, which DeepSeek identifies as V4.1 Flash. The older deepseek-v4-flash alias remains accepted for compatibility, but routes to V4.1 Flash and uses the current Flash pricing. deepseek-v4-pro remains separately priced.
Both models are documented with a 1 million-token context window and maximum output of 384,000 tokens. The calculator rejects an average request already beyond those published boundaries instead of returning a misleading cost for an impossible request shape.
Thinking tokens belong in the output budget
Thinking is enabled by default in the current model family unless it is disabled. Reasoning can be exposed separately for observability while still contributing to billed output usage.
For a production estimate, use the API's returned output-token usage rather than estimating cost from the visible answer alone. In Responses API usage, reasoning can be broken out through output_tokens_details.reasoning_tokens, while total output usage remains the billing input that matters here.
Before trusting the monthly total in production
Concurrency
DeepSeek currently documents account-level concurrency limits of 2,500 for deepseek-flash and 500 for deepseek-v4-pro. Exceeding the applicable limit can return HTTP 429 even when the token budget is acceptable.
User isolation
The optional user_id mechanism affects cache and scheduling isolation. DeepSeek says not to place private user information in that identifier, so use an internal non-sensitive identifier when the feature is needed.
Outside the estimate
Taxes, granted balance, credits, retries, surrounding services, hosting, network costs, account-specific terms, and future price changes are outside the token estimate.
The arithmetic runs locally in your browser. There is no prompt or API key field, and changing a workload value does not send it to DeepSeek.
DeepSeek sources used for this page
Pricing checked: September 22, 2026.Current rates, model names, cache behavior, token accounting, thinking behavior, and request limits were checked against DeepSeek's own documentation: models and pricing, context caching, token usage, thinking mode, rate limits and isolation. Recheck the provider pages before a launch or budget approval because pricing and limits can change.
