AI Cost Calculators
Gemini API Cost Calculator
Price Gemini API workloads across models, consumption modes, cache usage, grounding, and scheduled rate changes.
Build a Gemini workload estimate
Start with one representative request and monthly volume. Gemini pricing can then change with consumption mode, cache hits, input modality, prompt size, grounding, and—for some current Flash models—the pricing period.
Explicit cache storage and grounded searches
Open this when you create explicit cache objects or expect billable Google Search or Maps grounding.
Rates used in this estimate
Input$0.75 / 1M tokens
Cached input$0.075 / 1M tokens
Billed output$3.75 / 1M tokens
Explicit cache storage$0.50 / M token-hour
Search grounding (search queries)$14.00 / 1,000
Maps grounding$14.00 / 1,000
Worth checking before budgeting
- Google lists these Gemini 3.8/3.7/3.6 Flash rates through December 31, 2026 and higher built-in rates from January 1, 2027. The 12-month figure below simply repeats the selected rate; it is not a blended forward forecast.
Gemini cost breakdown
Token charges, explicit cache storage, and entered grounding usage are kept separate so you can see what is driving the estimate.
Estimated monthly cost
Per request
$0.006488
Per day
$12.975
12 months at selected rate
$4,671.00
Uncached input
270,000,000 tokens
$202.50
Cached input
90,000,000 tokens
$6.75
Billed output
48,000,000 tokens, including thinking when billed
$180.00
Requests: 60,000
Total input tokens: 360,000,000
Total billed output tokens: 48,000,000
* Important: Built-in Gemini Developer API rates were checked on September 21, 2026. Entered values are calculated locally in your browser; this calculator does not send prompts, workload figures, or API keys to Google. Verify current pricing and account-specific billing before committing spend.
Gemini pricing has more than one axis
A model name alone is not enough to reproduce a Gemini bill. The same request can use Standard, Batch, Flex, or Priority consumption, and those paths do not always share the same input, cached-input, output, or cache-storage rate. Some models also charge audio input differently from text, image, and video input.
Gemini 3.1 Pro Preview and Gemini 2.5 Pro add another boundary: prompts above 200,000 input tokens use the higher published token tier. The calculator derives that tier from the entered input tokens instead of asking for a second setting that could contradict the workload.
Batch and Flex change both price and workflow
Lower-cost processing often reduces input and output prices, but cached-input pricing can follow a different rule. Gemini 3.1 Pro and Gemini 2.5 Pro, for example, keep their Standard cache-read rate under Batch and Flex. Gemini 3.5 Flash also has a slightly different cached-input rate for Flex than Batch.
Batch is asynchronous and is designed around a turnaround of up to 24 hours. Flex stays synchronous, but runs on best-effort capacity with a 1–15 minute latency target. Those differences can matter as much as the token rate when deciding which path fits the workload.
Priority is not simply a price multiplier
Google publishes explicit Priority rates. If Priority capacity is exceeded, Google documents that requests can be downgraded to Standard and billed at the Standard rate. A budget based on Priority alone should therefore be compared with observed service-tier headers in production.
The 2026 introductory Flash price has an expiry date
Google currently lists introductory paid-tier pricing for Gemini 3.8 Flash, 3.7 Flash, and 3.6 Flash through December 31, 2026, with higher token and explicit-cache-storage rates from January 1, 2027. A normal “monthly cost × 12” label can therefore be misread as a forward-year forecast.
The calculator therefore lets you choose the published pricing period and labels the annualized number as “12 months at selected rate.” It does not blend the remaining 2026 months with 2027 pricing because no project start month is being requested.
Do not carry the introductory rate into a 2027 budget
For a deployment that will run into 2027, compare both pricing periods or build a month-by-month forecast. The current-period annualized figure is deliberately not presented as a forecast.
Cache hits and cache storage are separate charges
Gemini supports implicit and explicit context caching. Implicit caching is automatic on supported newer models, but Google does not guarantee a cache hit. When real usage data is available, the cached-token count is a better basis than assuming that a fixed percentage of every prompt will receive the discount.
Explicit caching adds a second billing dimension: stored tokens are charged for the time they remain cached. The optional “million token-hours” input represents the aggregate storage footprint directly, so a cache that is created, replaced, or held for different TTLs can still be represented without pretending every cached token stays stored for a full month.
The Interactions API and GenerateContent API also differ here: Google documents explicit cache objects for GenerateContent, while the Interactions API uses implicit caching. Leave explicit cache storage at zero when you are only modelling implicit cache hits.
Grounding is billed in its own units
Google Search and Google Maps grounding should not be converted into token charges. Gemini 3.x pricing uses a shared monthly allowance and then charges for billable search activity; Google also notes that one customer request can trigger more than one Search query. Gemini 2.5 pricing uses different free allowances and per-1,000 charges.
Because those allowances can be shared across models or depend on daily usage, the calculator does not guess how much free grounding remains in your account. Enter only the Search or Maps units that you expect to be billable after the applicable allowance.
Maps availability depends on the consumption path
Current Gemini 2.5 pricing does not list Google Maps grounding as available under Batch or Flex for the Pro, Flash, and Flash-Lite models covered here. The Maps field disappears for those combinations instead of calculating a charge for an unavailable path.
Billed output can be larger than the answer on screen
Google states that the listed output rate includes thinking tokens. With thinking enabled, billing is based on generated answer tokens plus thought tokens, even though the full internal reasoning is not returned as ordinary response text. Enter billed output usage rather than estimating cost from the visible answer alone.
For measured workloads, Gemini usage metadata exposes input, output, cached-content, and thinking token counts. Google's token-counting endpoint can also measure request input before a generation call. Those values are safer for production budgets than converting words or characters into tokens by hand.
Boundaries of this estimate
The model list is intentionally limited to text-output Gemini models that fit this token-based estimator. Live audio models, TTS, transcription, native image generation, video generation, embeddings, and other products can use different units or output pricing and should not be forced into the same calculation.
The estimate also cannot know your remaining free-tier or grounding allowance, retries that were not included in request volume, negotiated terms, taxes, currency conversion, regional or account-specific conditions, or future pricing changes beyond the rates Google has already published. Custom rates are available for cases where your contract or a newer pricing page differs from the built-in values.
Workload values stay in the browser. No Gemini request is made, and there is no reason to paste an API key, prompt text, customer data, or cached content into the calculator.
Official Google sources
Built-in rates and billing behavior were checked on September 21, 2026. These are the Google pages used to verify the model rates, cache treatment, consumption modes, token accounting, and availability rules represented above.
