Calculator

A/B model routing cost calculator

Compare all-cheap, all-premium, and a routed blend. Use a fixed A/B split or a cascade (try the cheap model first, escalate a share to premium).

Routed blend

Blended monthly cost

$1,025.00$0.00205 / req

Save $2,725.00 (72.7%) vs all-premium GPT-5

All GPT-5 nano$150.00
$0.0003 / req
Routed blend$1,025.00
$0.00205 / req
All GPT-5$3,750.00
$0.0075 / req
A-only requests400,000
B requests100,000
Premium avoided80.0%
Vs all-cheap$875.00

Routing

Models & traffic split

Every request hits A; a share also runs on B.

Playbook

Savings vs all-premium

Traffic mix drives blended spend.

Savings vs all-premium72.7% · $2,725.00
Premium traffic avoided80.0%
A cost / request$0.0003
B cost / request$0.0075
Extra A attempts (escalate)100,000
Blended monthly$1,025.00

Vs all-cheap: spend $875.00 more for quality headroom on escalated / B traffic.

Routing quality gates are product-specific—this prices the traffic mix, not evaluator accuracy. Updated 2026-07-31. Approximate static list prices for planning only. Always verify against each provider’s official pricing page before production budgeting.

Guide

How to get value from this calculator

Price A/B splits and cheap-first cascade routing between two models to find blended cost at target quality. Routing a fraction of traffic to a smaller model cuts average cost if the cheap path handles most requests successfully. Model cascade economics before building a router.

Cost formula

Blended cost = (traffic share A × cost per request model A) + (traffic share B × cost per request model B); Cascade cost = cheap model cost + (escalation rate × premium model cost)

Why it matters

Intelligent routing is the highest-leverage cost optimization for multi-task products, but blind 50/50 A/B tests waste quality or savings. Explicit cost modeling guides split ratios and escalation thresholds.

How to use it

  1. Define model A (cheap) and model B (premium) with token assumptions for each.
  2. Set traffic split percentage for A/B routing or escalation rate for cascade.
  3. Enter monthly request volume.
  4. Read the traffic-split playbook for blended cost vs 100% premium routing.
  5. Estimate quality risk of cheap-path failures and escalation frequency.
  6. Validate escalation rate in shadow mode before production routing.

Planning tips

  • Route classification and triage to the cheap model; reserve premium for complex generation.
  • Treat the traffic-split playbook as the decision receipt before you ship a router.
  • Cascade (try cheap first, escalate on low confidence) beats fixed splits for many workloads.
  • Log escalation reasons to tune thresholds without increasing premium traffic unnecessarily.
  • Output-heavy premium responses dominate blended cost—watch escalation on verbose tasks.
  • Prompt caching on shared prefixes benefits both tiers—include in per-model assumptions.

Frequently asked questions

What escalation rate should I plan for?

Start with 10–20% escalation in shadow testing if unsure. Production rates depend heavily on task difficulty and classifier quality.

Are model prices exact?

CentsPerToken uses approximate list prices for planning. Verify both models' rates officially.

Does routing add latency cost?

Cascade adds latency on escalated requests. This calculator focuses on token cost—factor latency in the latency-cost tool separately.

Can I route more than two models?

This tool compares two models. Extend the blended cost formula manually for multi-tier routers using the same per-model costs.