AI Cost Calculators
AI Reranking Cost Calculator
Compare current Voyage AI and Pinecone reranking costs, then estimate downstream LLM context savings and the net monthly impact.
Enter Your Reranking Workload
Compare provider billing and estimate the context cost removed before the LLM request.
Workload used for comparison
Successful searches: 100,000
Billable rerank requests: 105,000
Processed tokens per request: 21,000
Monthly processed tokens: 2,205,000,000
Context tokens before reranking: 2,000,000,000
Context tokens after reranking: 200,000,000
Reranking Cost and LLM Savings
The selected reranker is shown first, followed by a ranked paid list-price comparison.
Selected reranker monthly cost
Per successful search
$0.000441
Per 1,000 searches
$0.441
Per year
$529.20
Lowest cost · Voyage AI
rerank-2.5-lite
Paid usage after any account allowance
Billing rate
$0.02 /1M tokens
Per 1K searches
$0.441
Monthly
$44.10
Voyage AI
rerank-2.5
Paid usage after any account allowance
Billing rate
$0.05 /1M tokens
Per 1K searches
$1.1025
Monthly
$110.25
Pinecone
bge-reranker-v2-m3
Hosted inference paid list price
Billing rate
$2.00 /1K requests
Per 1K searches
$2.10
Monthly
$210.00
Pinecone
pinecone-rerank-v0
Hosted inference paid list price
Billing rate
$2.00 /1K requests
Per 1K searches
$2.10
Monthly
$210.00
Pinecone
cohere-rerank-v3.5
Cohere model hosted by Pinecone
Billing rate
$2.00 /1K requests
Per 1K searches
$2.10
Monthly
$210.00
Selected model: Voyage AI — rerank-2.5-lite
Context tokens removed before the LLM: 1,800,000,000
Estimated monthly LLM input saving: Enter the LLM input price
Net monthly impact after reranker cost: Enter the LLM input price
Possible monthly reranker saving against selected model: $0.00
Reranker budget status: $205.90 remaining
* Important: Calculated from the values currently shown. Default usage values are examples, so change them to match your expected usage. Built-in Voyage AI and Pinecone hosted reranking rates were checked on June 20, 2026. Final charges may include free allowances, embeddings, vector database charges, LLM output tokens, caching, hosting, engineering, taxes, discounts, and the business value of retrieval-quality changes.
Reranking adds another paid step to a RAG or search pipeline, but it can also reduce the amount of retrieved context sent to the language model. This calculator compares the reranker bill and the possible LLM input-token saving in one place.
Comparing Reranking Costs Across Billing Models
Enter monthly searches, candidate documents, average query length, average document length, and retry overhead. The calculator converts the same workload into each provider's official billing unit.
Voyage AI is calculated from processed tokens. Pinecone hosted rerankers are calculated from billable requests. A custom option supports either per-token or per-request pricing.
Results include monthly cost, cost per successful query, cost per 1,000 queries, annual cost, and a ranked comparison.
Estimating Downstream LLM Token Savings
Set how many candidate documents are retrieved and how many documents remain after reranking. The difference estimates the context tokens removed before the LLM request.
Enter the current LLM input price per one million tokens to estimate the monthly input-token saving. The tool subtracts the selected reranker cost to show a net monthly impact.
The estimate assumes one final LLM request per successful search. Multi-step agents, repeated prompts, caching, and provider-specific token rules can change the real result.
Choosing Candidate and Final Document Counts
Candidate documents are the initial search results sent to the reranker. Final documents are the highest-ranked results passed into the LLM prompt.
A larger candidate set can improve the chance of finding relevant information, but it increases token-based reranker usage. Sending fewer final documents can reduce LLM context cost, but removing too much context may reduce answer quality.
Practical Decisions This Tool Supports
- Compare token-based and request-based reranker pricing.
- Estimate reranking cost per query and per 1,000 queries.
- See how candidate-document depth changes the bill.
- Estimate LLM context tokens removed by reranking.
- Check whether token savings can offset reranker cost.
- Compare a public list price with a private provider quote.
- Plan annual reranking spend before production scale.
Built-In Providers and Pricing
The built-in comparison includes current Voyage AI rerank-2.5 and rerank-2.5-lite token prices, plus Pinecone hosted bge-reranker-v2-m3, pinecone-rerank-v0, and cohere-rerank-v3.5 request prices.
Direct Cohere pay-as-you-go pricing is not hardcoded because the current public pricing page defines a search unit but does not expose a clear direct production amount in the page content used for this check. Cohere direct pricing can still be entered through the custom option.
Prices were checked on June 20, 2026. Free monthly allowances and promotional credits are not deducted.
Costs and Benefits Not Included
The result does not automatically include embeddings, vector storage, vector reads, LLM output tokens, caching, hosting, observability, engineering work, taxes, support, latency, or the business value of better retrieval quality.
Use the RAG Cost Calculator for the wider retrieval stack and the AI Embedding Cost Comparison Calculator for document and query embedding costs.
Frequently Asked Questions
Why do Voyage AI and Pinecone use different billing calculations?
Voyage AI prices its current reranker endpoint by processed tokens. Pinecone prices its hosted reranking models by requests. The calculator applies the correct billing method to each option before comparing monthly cost.
How are Voyage reranker tokens calculated?
Voyage defines processed tokens as query tokens multiplied by the number of documents, plus the tokens in all documents. The calculator applies that formula to every billable reranking request.
Why does the tool estimate LLM input savings?
Reranking can reduce the number of retrieved documents sent to the language model. Fewer context documents can lower LLM input-token usage, although actual savings depend on the final prompt and workflow.
Does a positive net saving mean the reranker is automatically the best choice?
No. The estimate covers reranker cost and downstream LLM input-token savings. Retrieval quality, latency, relevance, conversion, support, limits, and engineering work also matter.
Why are free allowances excluded?
Free allowances can differ by plan and may not continue at production scale. The built-in comparison uses recurring paid list prices so the options are compared consistently.
Can I compare Cohere direct pricing or another provider?
Yes. Enable custom pricing and enter the current rate and billing unit from the provider's official page or your account quote.
