APIs win on low or spiky traffic
If GPUs sit idle most of the month, rental or owned hardware still burns fixed cost. Hosted APIs scale to zero more cleanly.
Local wins on steady, high volume
Once you keep accelerators busy and your token mix is predictable, GPU-hours can undercut flagship API list prices—especially without batch discounts.
Price ops, not just silicon
Add utilization, engineering time, and reliability risk. A spreadsheet that ignores on-call will overstate local savings.
FAQ
Is tokens/second a real benchmark?
No—use it as an editable planning input. Replace it with your measured end-to-end throughput for decisions.