decision · section 04 of 10

If you build with AI APIs

Which model, provider, and pricing structure should this workload commit to?

Measure effective cost per successful task rather than per token, then take the caching and batch discounts, then route — in that order.

Cache economics reverse for workloads without enough context reuse to amortize the cache-write premium.

The section itself

Don’t compare token prices — compare cost-per-quality-point for your workload. A model that is 10x cheaper per token but takes 3 retries to get a correct answer is more expensive than the one it replaced. The non-linear frontier means the biggest cost leverage is not picking the cheapest model, but picking the cheapest model that clears your quality bar — and that bar is workload-specific. Two caveats now carry the same weight as the rule itself: posted-price ranges are upper bounds (token-efficiency varies 2.65x+ by model), and the quality index behind any cost-per-quality ranking is an editorially- weighted composite — use it to find the non-linear region, not to rank models to a decimal.

Who uses this section, and for what

Reader questions that route through here

In sequence