Start from the task, not the brand
Classification and extraction often work on budget tiers. Hard reasoning and long agents usually need a higher capability floor. Match the workload before you chase the lowest sticker price.
Filter, then sort by cost
Set minimum capability, context window, and providers you can actually ship. Then rank by estimated API spend for your token sizes and request volume.
Validate with evals
The cheapest model on paper can still fail your product. Run a small offline eval or shadow traffic before you switch production routing.
FAQ
Does cheapest always mean best ROI?
No. Retries, human review, and worse answers can erase token savings. Pair cost ranking with quality checks.