Calculator

Cheapest model for this task

Choose a workload preset, set a minimum capability band, and see which models are cheapest for your token volume—before you overpay for a flagship tier you do not need.

Cheapest match

Grok 4 Fast

$37.30/ 100,000 requests

$0.000373 / req · xAI (Grok) · capable

Capabilitycapable
Context2,000,000
Vs most expensive$31,762.70 (100%)
Models considered45

Task

Describe the workload

Helpful customer replies with moderate context.

Podium

Ranked by monthly cost

View pricing

Runner-up: Grok 4.1 Fast Reasoning · $37.30

1
Grok 4 Fast
xAI (Grok) · capable · $0.000373 / req
$37.30
2
Grok 4.1 Fast Reasoning
xAI (Grok) · capable · $0.000373 / req
$37.30
3
Grok 4.1 Fast Non-Reasoning
xAI (Grok) · capable · $0.000373 / req
$37.30
4
DeepSeek Chat (V3.2)
DeepSeek · balanced · $0.000379 / req
$37.88
5
DeepSeek Reasoner (V3.2)
DeepSeek · capable · $0.000379 / req
$37.88
6
DeepSeek V3.2 Speciale
DeepSeek · balanced · $0.000379 / req
$37.88
7
Grok Code Fast 1
xAI (Grok) · balanced · $0.000723 / req
$72.30
8
GPT-3.5 Turbo
OpenAI · balanced · $0.000885 / req
$88.50

Recommends the lowest list-price model that passes your filters—not the best quality model. Always validate quality on real evals. Updated 2026-07-31. Approximate static list prices for planning only. Always verify against each provider’s official pricing page before production budgeting.

Guide

How to get value from this calculator

Find the lowest-cost model that still meets your capability floor for a given task. Enter input and output token volumes and minimum requirements to filter out models that are cheap but unsuitable. Saves money without sacrificing quality by matching task complexity to model tier.

Cost formula

Model total cost = (input tokens / 1M × input price) + (output tokens / 1M × output price); Cheapest fit = minimum total cost among models meeting capability floor

Why it matters

Teams often default to flagship models for tasks that smaller models handle well, paying 10–50x more than necessary. Systematic cheapest-fit selection aligns model capability with task difficulty and budget.

How to use it

  1. Describe your task type and set minimum capability requirements (context length, tools, vision, etc.).
  2. Enter expected input and output tokens per request or per month.
  3. Review the ranked podium of cheapest eligible models for your workload.
  4. Tighten the capability filter so unfit models drop out even if they look cheapest.
  5. Run quality evals on the top two cheapest eligible models before switching.
  6. Verify current pricing on the provider's site—cheap today may not be cheap next quarter.

Planning tips

  • Classification, extraction, and routing tasks often work on the smallest model in a family.
  • Use the capability filter before trusting the podium—cheap models that fail your floor are excluded.
  • Output token pricing drives cost for verbose tasks—cheapest input rate does not guarantee cheapest total.
  • Consider batch pricing if latency requirements allow—it can reorder the cheapest-fit ranking.
  • A slightly more expensive model with prompt caching may beat a cheaper model without cache on repeated prefixes.
  • Treat the podium as a shortlist—run evals on the top finishers before switching.
  • Re-run this analysis when providers release new small models or cut prices.

Frequently asked questions

Does cheapest mean lowest quality?

Not necessarily—it means lowest cost among models you marked as capable for the task. Always validate quality with evals before production deployment.

How are capability floors defined?

You specify requirements like minimum context window, tool support, or modality. Models that fail any requirement are excluded from the cheapest ranking.

Are prices official?

CentsPerToken uses approximate list prices for planning. Confirm rates officially and factor in any enterprise discounts you have.

Should I optimize for per-request or monthly cost?

Monthly cost = per-request cost × volume. High-volume workloads should rank on monthly total, not per-request input rate alone.