AI Cost Calculators
Grok API Cost Calculator
Estimate a Grok workload using the xAI pricing rules that actually change the bill: context length, cache hits, service tier, regional processing, and server-side tools.
Shape the Grok request you expect to run
Start with the model and service path, then use average billed token counts from one representative request.
Short-context rates (prompt < 200,000 tokens)
xAI applies the active context band to uncached input, cached input, and output tokens for the request.
Grok cost breakdown
Token charges and any server-side tool usage are kept separate so you can see what is driving the estimate.
Estimated monthly cost
Per request
$0.004140
Token charges
$289.80
Tool charges
$0.00
Uncached input
67,200,000 tokens
$134.40
Cached input
16,800,000 tokens
$8.40
Billed output
24,500,000 tokens
$147.00
Model: grok-4.7
Pricing band: short context
Processing: Standard · Global endpoint
Effective / 1M tokens: input $2.00, cached $0.50, output $6.00
12 months at this monthly workload: $3,477.60
* Important: Calculated from the values currently shown. Default usage values are examples, so change them to match your expected usage. Built-in xAI rates were checked on September 22, 2026. Final charges may include image and video generation, voice APIs, file or collection storage and downloads, client-side tool costs, taxes, credits, negotiated discounts, retries not represented in the workload, and any usage not entered here.
The 200K prompt boundary can double the token rate
xAI publishes separate short- and long-context prices for the current Grok text models. Once a request reaches the 200,000-token prompt threshold, the long-context rate applies to all input, cached input, and output tokens in that request—not only the tokens above the boundary.
That makes average prompt size more important than a monthly token total alone. Two workloads can consume the same number of tokens in a month while landing in different pricing bands because one sends many smaller prompts and the other sends fewer very large prompts.
Cached tokens are visible in the API response
xAI automatically caches matching prompt prefixes. A planning percentage is useful before launch, but production budgets should come from the returned cached_tokens value. If cache hits stay at zero across a continuing conversation, xAI recommends checking the conversation or prompt-cache key and whether earlier messages are changing.
Reasoning tokens are different: they are billed at the output-token rate. For reasoning models, use billed output usage rather than estimating cost from visible answer length alone.
Priority and Batch solve different problems
Priority Processing is for latency-sensitive real-time requests and costs 2× the standard token rates when the response confirms the priority tier. Batch is asynchronous, normally completes within 24 hours, and currently gives a 20% token discount only on Grok 4.3 and the Grok 4.20 variants listed by xAI.
Agentic Grok requests can spend outside the token line item
Web Search, X Search, code execution, attachment search, and collection search have their own invocation charges. The model can also make more than one server-side tool call while answering a single user request, so request count is not a safe substitute for tool usage. xAI exposes successful billable usage separately; that is the number to use when you have production data.
X Search changed on September 21, 2026: posts fetched are billed at $5 per 1,000 and user profiles at $10 per 1,000. The calculator therefore asks for those fetched-item counts rather than pretending X Search still has one flat per-call price.
What this estimate deliberately keeps separate
Image and video generation, voice APIs, xAI file and collection storage, download charges, client-side tools, taxes, credits, negotiated pricing, and retries that are not already represented in the entered workload are outside this estimate. Multi-agent work also needs aggregate billed token usage because leader and sub-agent activity is chargeable even when only the leader's final answer is returned.
The arithmetic runs in your browser. Beeija does not send the workload values entered here to xAI. Following an official documentation link opens xAI's site separately.
xAI documentation behind the calculation
Built-in rates were checked on September 22, 2026 against xAI's API pricing and model reference. Cache accounting follows the prompt-caching usage guide. The processing choices come from the Priority Processing and Batch API documentation, while optional tool costs follow xAI's current server-side tool pricing.
