What this section answers
Measure effective cost per successful task rather than per token, then take the caching and batch discounts, then route — in that order.
decision · section 04 of 10
Which model, provider, and pricing structure should this workload commit to?
What this section answers
Measure effective cost per successful task rather than per token, then take the caching and batch discounts, then route — in that order.
Boundary
Cache economics reverse for workloads without enough context reuse to amortize the cache-write premium.
Source coordinate
Report heading If you build with AI APIs, under What this means for decisions, in The real cost of AI: August 2026.
Don’t compare token prices — compare cost-per-quality-point for your workload. A model that is 10x cheaper per token but takes 3 retries to get a correct answer is more expensive than the one it replaced. The non-linear frontier means the biggest cost leverage is not picking the cheapest model, but picking the cheapest model that clears your quality bar — and that bar is workload-specific. Two caveats now carry the same weight as the rule itself: posted-price ranges are upper bounds (token-efficiency varies 2.65x+ by model), and the quality index behind any cost-per-quality ranking is an editorially- weighted composite — use it to find the non-linear region, not to rank models to a decimal.
The decision as a sequence: measure effective cost, take the discounts, then route.
The levers available before you change the product.