AI Cost Calculators
Mistral API Cost Calculator
Budget Mistral text inference using the service tier, region, cache behavior, and workload shape that will actually be deployed.
Plan a Mistral inference workload
Choose how the request is served, then enter one representative request and monthly volume. Cache hits are priced separately from ordinary input.
Mistral Small 4 is represented by the API alias mistral-small-latest and uses a 256,000-token context window here. The limit check applies to the representative request entered above; production code still needs to validate each actual prompt plus its allowed generation budget.
The arithmetic runs in your browser. These planning values are not sent to Mistral, and no API key or prompt text is required.
Mistral cost breakdown
Uncached input, cached input, and output are kept separate so the effect of prompt caching stays visible.
Estimated monthly cost
Per request
$0.000345
Daily avg. (30d)
$0.6906
12-mo projection
$248.616
Uncached input
52,800,000 tokens
$7.92
Cached input
13,200,000 tokens
$0.198
Output
21,000,000 tokens
$12.60
Model: Mistral Small 4 · standard · global
Input tokens: 66,000,000 · cached 13,200,000
Output tokens: 21,000,000
12-month projection assumes the selected rates stay unchanged.
* Important: Calculated from the values currently shown. Default usage values are examples, so change them to match your expected usage. Built-in Mistral rates were checked on September 24, 2026. Final charges may include OCR, transcription, text-to-speech, fine-tuning, agents, built-in tools, files, storage, taxes, credits, negotiated terms, retries, and services billed in units other than text tokens.
Mistral's service tiers change more than the token rate
Standard is the ordinary synchronous path. Batch is for work that can wait in an asynchronous queue and is published at a 50% discount. Priority is intended for real-time or business-critical traffic, costs 1.75× Standard list pricing, and requires Mistral to configure Priority capacity for the organization.
| Path | Cost treatment | Operational trade-off |
|---|---|---|
| Standard | Published list rate | Best-effort synchronous processing |
| Batch | 50% lower | Asynchronous; queued for processing over a 24-hour period |
| Priority | 1.75× Standard | Priority queue; can fall back to Standard when configured Priority capacity is unavailable |
Cache-hit input is a separate billing path
Mistral prompt caching reuses a compatible prompt prefix. Cached prompt tokens are billed at 10% of the Standard input rate, which is why the calculator asks for a cache-hit share instead of applying one input price to every prompt token. The API exposes measured cache usage in usage.prompt_tokens_details.cached_tokens; ordinary billable input is the remaining prompt-token count.
A cache key can improve the chance of reuse, but it does not guarantee a hit. Mistral also notes that cache blocks contain 64 tokens, so very short prompts will not produce cache hits. For an existing application, measured cached-token usage is a better budgeting input than a guessed percentage.
Regional inference is a deployment decision, not just a 10% surcharge
Mistral offers Global, EU, and US inference endpoints. EU and US regional inference add 10% to input, cached-input, and output token pricing, but they also change where eligible inference is processed. That can matter for data-location requirements and latency.
Regional processing does not make the entire Mistral control plane regional. Account settings, API keys, billing, access management, analytics, and other operational metadata can still be handled outside the selected inference geography. Regional inference and Zero Data Retention are also separate controls: one governs where eligible inference runs, while the other governs whether eligible request and response content is retained after processing.
Regional endpoints also narrow feature choices
Batch, Agents, and the Files API are not available on regional endpoints, and model availability varies by region. Function calling is currently the supported regional tool path. That is why choosing Batch in the calculator returns the inference location to Global instead of pretending those options can be combined.
Reasoning can increase the output side of the bill
Mistral Small 4 and Mistral Medium 3.5 support adjustable reasoning_effort. Higher reasoning can generate a thinking chunk before the final answer and uses more generated tokens. When reasoning is enabled, budget from measured completion usage rather than counting only the final visible answer.
Context limits are a hard request boundary
Mistral counts both input and generated output tokens toward each request's context limit. Requests that exceed the model limit return a 400 error. The calculator checks the representative input-plus-output request entered above, but an average cannot guarantee that every production request fits; validate each real prompt together with its allowed generation budget.
256,000 tokens: Mistral Large 3, Medium 3.5, Small 4, and Ministral 3
128,000 tokens: Codestral
A -latest model ID can move underneath a budget
Mistral's -latest aliases automatically move to newer General Availability versions. That is convenient during development, but Mistral warns that an alias can expose an application to changes in model behavior and pricing. For a production budget that must stay reproducible, pin the specific major.minor model identifier you actually intend to deploy and revisit the estimate when you migrate.
What this text-token estimate intentionally leaves out
The calculation stops at hosted text-token inference. OCR, transcription, text-to-speech, fine-tuning, Agents, built-in tools, Files, storage, retries, taxes, credits, negotiated terms, and products billed by pages, minutes, characters, or another unit need their own cost model. Keeping those units separate is more useful than hiding them inside a generic miscellaneous charge.
Mistral documentation behind the calculation
Built-in rates were checked on September 24, 2026. Current model prices and cached-input rates come from Mistral's pricing documentation. Delivery behavior and the 1.75× premium follow the Priority Tier documentation, while the 50% asynchronous discount follows the Batch processing documentation. Regional limits and the 10% surcharge are described in the regional inference documentation, cache-hit accounting follows Mistral's prompt caching documentation, reasoning behavior follows the reasoning documentation, and alias stability follows the model lifecycle policy.
