Multi-model API cost calculator
Estimate input/output spend across OpenAI, Anthropic, Gemini, Perplexity, DeepSeek, and Grok.
CentsPerToken tools directory
Estimate costs for OpenAI, Anthropic, Gemini, Grok, DeepSeek, Mistral, Perplexity and more—all in one place.
Featured
Start with the most-used CentsPerToken tools.
Estimate input/output spend across OpenAI, Anthropic, Gemini, Perplexity, DeepSeek, and Grok.
Estimate tokens before you send a prompt or batch.
Scan list prices side by side across major providers.
Model indexing, retrieval, and generation spend for RAG apps.
Turn product traffic assumptions into a monthly AI budget.
Price multi-step agents with tool rounds, growing context, and cache.
Full catalog
33 tools
Estimate input/output spend across OpenAI, Anthropic, Gemini, Perplexity, DeepSeek, and Grok.
Estimate tokens before you send a prompt or batch.
Scan list prices side by side across major providers.
Model indexing, retrieval, and generation spend for RAG apps.
See how much 1M tokens costs by model—and how many tokens your budget buys.
Turn product traffic assumptions into a monthly AI budget.
Price multi-step agents with tool rounds, growing context, and cache.
Estimate savings from prompt caching with hit rate and cache writes.
Price image runs for DALL·E, Flux, and similar APIs.
Recommend the lowest-cost model that fits your task and capability floor.
See how much batch pricing can cut from realtime traffic.
Price 8K–1M context packs, check fit, and find cheaper models that still fit.
Estimate page→token, chunked embedding, and LLM extract cost for document corpora.
Map API COGS, adoption, and seat/usage pricing to margin and price floors.
Price A/B splits or cheap-first cascade routing between two models.
Compare provider API cost to what you charge customers—margin and price floors.
Compare RAG with vs without LLM rerank—when top-N context savings beat rerank spend.
Estimate savings from shorter prompts and cached context.
Price candidate generation and pointwise/pairwise judge scoring for eval suites.
See which RPM/TPM/daily or concurrency limit binds—and cost at capacity.
Estimate generate → critique → filter cost for synthetic datasets, including over-generation.
Compare chat-style usage against completion-style workloads.
Rank models by API cost and planning latency under your SLO budget.
Compare hosted API spend against rented or owned GPU-hour economics.
Split monthly AI budget across seats and roles with per-seat request caps.
Project training and inference costs for fine-tuned models.
Paste usage exports to estimate costs and compare with invoice totals locally.
Compare early-cancel savings, TTFT, and wait economics for streamed LLM responses.
Convert USD AI API bills into GBP, EUR, AUD, and CAD for local budgeting.
Budget Whisper, TTS, and other audio endpoints.
Estimate spend when a primary model fails over to a secondary provider.
Estimate how retries, timeouts, and failure billing inflate LLM API spend.
Rough cost planning for video and multimodal workloads.
Browse by category
Scan the catalog by workload instead of scrolling one long list.
Core token and API spend estimators for production workloads.
Estimate input/output spend across OpenAI, Anthropic, Gemini, Perplexity, DeepSeek, and Grok.
Price multi-step agents with tool rounds, growing context, and cache.
See how much batch pricing can cut from realtime traffic.
Compression, caching, and context-window cost planning.
Estimate savings from prompt caching with hit rate and cache writes.
Price 8K–1M context packs, check fit, and find cheaper models that still fit.
Estimate savings from shorter prompts and cached context.
Price image runs across popular generation APIs.
Price image runs for DALL·E, Flux, and similar APIs.
Rough multimodal and video workload cost planning.
Rough cost planning for video and multimodal workloads.
Speech-to-text, TTS, and voice API budgeting.
Budget Whisper, TTS, and other audio endpoints.
RAG indexing, retrieval, rerank, and document ingestion.
Model indexing, retrieval, and generation spend for RAG apps.
Estimate page→token, chunked embedding, and LLM extract cost for document corpora.
Compare RAG with vs without LLM rerank—when top-N context savings beat rerank spend.
Training, synthetic data, and eval/judge suite costs.
Price candidate generation and pointwise/pairwise judge scoring for eval suites.
Estimate generate → critique → filter cost for synthetic datasets, including over-generation.
Project training and inference costs for fine-tuned models.
Tokenizers, CSV analyzers, rate limits, and streaming math.
Estimate tokens before you send a prompt or batch.
Compare early-cancel savings, TTFT, and wait economics for streamed LLM responses.
Monthly budgets, seats, margins, and SaaS unit economics.
Turn product traffic assumptions into a monthly AI budget.
Map API COGS, adoption, and seat/usage pricing to margin and price floors.
Compare provider API cost to what you charge customers—margin and price floors.
Latency, retries, failover, and usage export analysis.
Rank models by API cost and planning latency under your SLO budget.
Paste usage exports to estimate costs and compare with invoice totals locally.
Estimate how retries, timeouts, and failure billing inflate LLM API spend.
Model routing, cheapest-fit, chat vs completion, local vs cloud.
Recommend the lowest-cost model that fits your task and capability floor.
Price A/B splits or cheap-first cascade routing between two models.
Compare chat-style usage against completion-style workloads.
Local GPUs, concurrency, and capacity planning.
See which RPM/TPM/daily or concurrency limit binds—and cost at capacity.
Compare hosted API spend against rented or owned GPU-hour economics.
Estimate spend when a primary model fails over to a secondary provider.
List-price tables, 1M-token costs, and currency conversion.
Scan list prices side by side across major providers.
See how much 1M tokens costs by model—and how many tokens your budget buys.
Convert USD AI API bills into GBP, EUR, AUD, and CAD for local budgeting.
Why CentsPerToken
Planning rates stamped with last-updated dates across the catalog.
Most tools respond in under a second entirely in-browser.
Open any calculator and start estimating immediately.
New models and cost scenarios land without a migration project.
Static pages and light client filters keep browsing snappy.
Core CentsPerToken tools stay free with no account wall.