Cost formula
Blended cost = (traffic share A × cost per request model A) + (traffic share B × cost per request model B); Cascade cost = cheap model cost + (escalation rate × premium model cost)
Calculator
Compare all-cheap, all-premium, and a routed blend. Use a fixed A/B split or a cascade (try the cheap model first, escalate a share to premium).
Routed blend
$1,025.00$0.00205 / req
Save $2,725.00 (72.7%) vs all-premium GPT-5
Routing
Every request hits A; a share also runs on B.
Playbook
Traffic mix drives blended spend.
Vs all-cheap: spend $875.00 more for quality headroom on escalated / B traffic.
Routing quality gates are product-specific—this prices the traffic mix, not evaluator accuracy. Updated 2026-07-31. Approximate static list prices for planning only. Always verify against each provider’s official pricing page before production budgeting.
Guide
Price A/B splits and cheap-first cascade routing between two models to find blended cost at target quality. Routing a fraction of traffic to a smaller model cuts average cost if the cheap path handles most requests successfully. Model cascade economics before building a router.
Blended cost = (traffic share A × cost per request model A) + (traffic share B × cost per request model B); Cascade cost = cheap model cost + (escalation rate × premium model cost)
Intelligent routing is the highest-leverage cost optimization for multi-task products, but blind 50/50 A/B tests waste quality or savings. Explicit cost modeling guides split ratios and escalation thresholds.
Start with 10–20% escalation in shadow testing if unsure. Production rates depend heavily on task difficulty and classifier quality.
CentsPerToken uses approximate list prices for planning. Verify both models' rates officially.
Cascade adds latency on escalated requests. This calculator focuses on token cost—factor latency in the latency-cost tool separately.
This tool compares two models. Extend the blended cost formula manually for multi-tier routers using the same per-model costs.