Beeija
← Back to Tools

AI Cost Calculators

AI Reranking Cost Calculator

Compare current Voyage AI and Pinecone reranking costs, then estimate downstream LLM context savings and the net monthly impact.

Enter Your Reranking Workload

Compare provider billing and estimate the context cost removed before the LLM request.

Workload used for comparison

Successful searches: 100,​000

Billable rerank requests: 105,​000

Processed tokens per request: 21,​000

Monthly processed tokens: 2,​205,​000,​000

Context tokens before reranking: 2,​000,​000,​000

Context tokens after reranking: 200,​000,​000

Reranking Cost and LLM Savings

The selected reranker is shown first, followed by a ranked paid list-price comparison.

Selected reranker monthly cost

$44.10

Per successful search

$0.000441

Per 1,000 searches

$0.441

Per year

$529.20

Lowest cost · Voyage AI

rerank-2.5-lite

Paid usage after any account allowance

Billing rate

$0.02 /1M tokens

Per 1K searches

$0.441

Monthly

$44.10

Voyage AI

rerank-2.5

Paid usage after any account allowance

Billing rate

$0.05 /1M tokens

Per 1K searches

$1.1025

Monthly

$110.25

Pinecone

bge-reranker-v2-m3

Hosted inference paid list price

Billing rate

$2.00 /1K requests

Per 1K searches

$2.10

Monthly

$210.00

Pinecone

pinecone-rerank-v0

Hosted inference paid list price

Billing rate

$2.00 /1K requests

Per 1K searches

$2.10

Monthly

$210.00

Pinecone

cohere-rerank-v3.5

Cohere model hosted by Pinecone

Billing rate

$2.00 /1K requests

Per 1K searches

$2.10

Monthly

$210.00

Selected model: Voyage AI — rerank-2.5-lite

Context tokens removed before the LLM: 1,​800,​000,​000

Estimated monthly LLM input saving: Enter the LLM input price

Net monthly impact after reranker cost: Enter the LLM input price

Possible monthly reranker saving against selected model: $0.00

Reranker budget status: $205.90 remaining

* Important: Calculated from the values currently shown. Default usage values are examples, so change them to match your expected usage. Built-in Voyage AI and Pinecone hosted reranking rates were checked on June 20, 2026. Final charges may include free allowances, embeddings, vector database charges, LLM output tokens, caching, hosting, engineering, taxes, discounts, and the business value of retrieval-quality changes.

Reranking adds another paid step to a RAG or search pipeline, but it can also reduce the amount of retrieved context sent to the language model. This calculator compares the reranker bill and the possible LLM input-token saving in one place.

Comparing Reranking Costs Across Billing Models

Enter monthly searches, candidate documents, average query length, average document length, and retry overhead. The calculator converts the same workload into each provider's official billing unit.

Voyage AI is calculated from processed tokens. Pinecone hosted rerankers are calculated from billable requests. A custom option supports either per-token or per-request pricing.

Results include monthly cost, cost per successful query, cost per 1,000 queries, annual cost, and a ranked comparison.

Estimating Downstream LLM Token Savings

Set how many candidate documents are retrieved and how many documents remain after reranking. The difference estimates the context tokens removed before the LLM request.

Enter the current LLM input price per one million tokens to estimate the monthly input-token saving. The tool subtracts the selected reranker cost to show a net monthly impact.

The estimate assumes one final LLM request per successful search. Multi-step agents, repeated prompts, caching, and provider-specific token rules can change the real result.

Choosing Candidate and Final Document Counts

Candidate documents are the initial search results sent to the reranker. Final documents are the highest-ranked results passed into the LLM prompt.

A larger candidate set can improve the chance of finding relevant information, but it increases token-based reranker usage. Sending fewer final documents can reduce LLM context cost, but removing too much context may reduce answer quality.

Practical Decisions This Tool Supports

  • Compare token-based and request-based reranker pricing.
  • Estimate reranking cost per query and per 1,000 queries.
  • See how candidate-document depth changes the bill.
  • Estimate LLM context tokens removed by reranking.
  • Check whether token savings can offset reranker cost.
  • Compare a public list price with a private provider quote.
  • Plan annual reranking spend before production scale.

Built-In Providers and Pricing

The built-in comparison includes current Voyage AI rerank-2.5 and rerank-2.5-lite token prices, plus Pinecone hosted bge-reranker-v2-m3, pinecone-rerank-v0, and cohere-rerank-v3.5 request prices.

Direct Cohere pay-as-you-go pricing is not hardcoded because the current public pricing page defines a search unit but does not expose a clear direct production amount in the page content used for this check. Cohere direct pricing can still be entered through the custom option.

Prices were checked on June 20, 2026. Free monthly allowances and promotional credits are not deducted.

Costs and Benefits Not Included

The result does not automatically include embeddings, vector storage, vector reads, LLM output tokens, caching, hosting, observability, engineering work, taxes, support, latency, or the business value of better retrieval quality.

Use the RAG Cost Calculator for the wider retrieval stack and the AI Embedding Cost Comparison Calculator for document and query embedding costs.

Frequently Asked Questions

Why do Voyage AI and Pinecone use different billing calculations?

Voyage AI prices its current reranker endpoint by processed tokens. Pinecone prices its hosted reranking models by requests. The calculator applies the correct billing method to each option before comparing monthly cost.

How are Voyage reranker tokens calculated?

Voyage defines processed tokens as query tokens multiplied by the number of documents, plus the tokens in all documents. The calculator applies that formula to every billable reranking request.

Why does the tool estimate LLM input savings?

Reranking can reduce the number of retrieved documents sent to the language model. Fewer context documents can lower LLM input-token usage, although actual savings depend on the final prompt and workflow.

Does a positive net saving mean the reranker is automatically the best choice?

No. The estimate covers reranker cost and downstream LLM input-token savings. Retrieval quality, latency, relevance, conversion, support, limits, and engineering work also matter.

Why are free allowances excluded?

Free allowances can differ by plan and may not continue at production scale. The built-in comparison uses recurring paid list prices so the options are compared consistently.

Can I compare Cohere direct pricing or another provider?

Yes. Enable custom pricing and enter the current rate and billing unit from the provider's official page or your account quote.

Explore Related AI Cost Tools