Set the multimodal workload — model, image count, resolution, traffic. The calculator shows how image tokens dominate the cost equation, and how the same workload's bill can vary 50× across vision-capable models. The visceral lesson: adding 5 high-res images to each request typically multiplies your bill by 5-10×. Most teams don't do this math before shipping vision features and discover the cost only when the bill arrives.