Set the workload — model, context length, output length, traffic. The calculator computes per-request and per-year cost using real 2026 hosted-API pricing, with toggles for prompt caching and the Batch API. The visceral lesson: small changes to model choice and caching strategy produce 10-50× cost differences. The team that ships at scale isn't the team with the best model — it's the team with the best routing. Use this to project your monthly bill before you build something at 1000 RPS only to discover it costs $5M/year.