Cost formula
Cloud cost = monthly tokens priced at API rates; Local cost = (GPU hourly rate × hours run) + power + ops overhead; Break-even = cloud monthly cost ÷ local monthly cost at given utilization
Calculator
Pick a cloud API model and a local or rented GPU profile. CentsPerToken charts both cost curves, marks the break-even request volume, folds engineering time into the fixed line, and checks whether your fleet can actually serve the traffic.
Verdict at your volume
$3,017.73saved / month
96% cheaper at 150,000 requests per month
Crossover analysis
Local never overtakes cloud with these settings—variable GPU cost per request is already higher than the API rate.
Step 1
Step 2
This volume needs 1242 GPU-hours but 1 GPU only supply 730. Add hardware or raise throughput.
Step 3
Self-hosting is rarely just GPU hours. Engineering time folds into the fixed monthly line so break-even stays honest.
29% of local spend is fixed overhead — the part that does not shrink when traffic drops.
Sensitivity
Local throughput and GPU rates are rough planning inputs—not MLPerf results. Capacity headroom assumes 730 hours per GPU per month. Self-hosting also adds reliability, security, and on-call risk not priced here. Updated 2026-07-31. Approximate static list prices for planning only. Always verify against each provider’s official pricing page before production budgeting.
Guide
Compare hosted LLM API spend against self-hosted or rented GPU economics for your workload. Cloud APIs win at low and variable volume; owned or dedicated GPUs can win at high stable throughput—but only after accounting for hardware, ops, and idle capacity. Use this for infrastructure decisions, not precise TCO audits.
Cloud cost = monthly tokens priced at API rates; Local cost = (GPU hourly rate × hours run) + power + ops overhead; Break-even = cloud monthly cost ÷ local monthly cost at given utilization
Local inference looks cheap on GPU-hour spreadsheets but loses when utilization is low or models update frequently. A structured comparison prevents premature infrastructure investment or overpaying cloud at scale.
Open-weight models may have license constraints. This calculator focuses on hardware and API cost comparison—verify license terms separately.
CentsPerToken uses approximate rental rates for planning. Your cloud GPU provider, spot pricing, and utilization will differ.
Be conservative. Production clusters rarely sustain 100% utilization. Model at 30–60% unless you have proven steady load.
No. API side uses approximate list prices. Confirm both API and GPU rates before major infrastructure decisions.