Proof standard
A price you can reproduce from a posted rate card, and a quality adjustment whose weighting you can inspect.
Audience route
You are choosing a model and a pricing structure for a workload you will operate.
The decision
Proof standard
A price you can reproduce from a posted rate card, and a quality adjustment whose weighting you can inspect.
Working output
An effective cost per successful task for your own workload — not a per-token rate.
The sections below are the report's own material, composed here. Nothing is rewritten for this audience; only the order and the framing are.
The four findings that change the answer before any per-model comparison starts.
Open this section on its ownThe sticker price of a token is not the cost of an answer. In August 2026, the same standard API workload (100K input + 20K output tokens) costs anywhere from $0.018 on Gemini Flash-Lite to $2.00 on Claude Fable 5 — a 111x price range across comparable models. Two corrections cut the other way: measured token-efficiency variance means the same input produces 2.65x+ more output tokens on some models, and cheap-token models can cost more per successful task via verbosity and retry loops — so the posted range is an upper bound on effective dispersion, and it still understates the real spread: adjusted for quality using the Artificial Analysis Intelligence Index, the cost-per-quality- point spreads by 47x. Read that 47x as order-of-magnitude evidence of non-linear pricing, not a precise ratio: the index is an editorially-weighted nine-benchmark composite (text/English-centric; agentic-weighted since v4.1), and version overhauls move scores by ~23 points — as large as the frontier spread itself. The marginal cost of capability accelerates sharply above Intelligence Index 50.
Running your own GPUs is a crossover problem, not a preference. Specialist GPU clouds (“neoclouds” like RunPod, Lambda, Together AI) charge 50-70% less than hyperscalers for the same H100: a median of $3.99/GPU-hour vs $7.89. The TCO model — now corrected for procurement scope — shows the answer splits by who’s buying: a lean operator adding GPUs to existing infrastructure (~$30K/GPU street price) breaks even against hyperscaler on-demand at ~21% utilization and beats median neocloud pricing above ~37%; an enterprise node-loaded buyer (~$94K/GPU all-in per io.net’s measured 3-year figure) breaks even vs hyperscaler on-demand only at ~57%, essentially never beats median neocloud pricing, and never beats cheap neoclouds or spot. Reserved commitments beat node-loaded on-prem at every utilization. Utilization assumption AND procurement scope together dominate the decision.
The control premium is real, and buyers pay it knowingly. Independent cost itemization puts true self-hosting at 3–5× the pure GPU price — matching the node-loaded math. For a compliance-driven subset, the premium is not optional: sovereign cloud IaaS spending is forecast at $80B in 2026 (+35.6%), Italy’s cloud market grew 20% YoY explicitly on sovereign-data demand, and the EU’s Technological Sovereignty Package (June 2026) extends data-governance obligations to AI providers. The other evidenced motives: guaranteed capacity with no rate limits or peak contention, latency for on-network workloads, vendor independence as a hedge against the list-price and subsidy moves forecast above, and very-high sustained utilization. In practice the control premium runs roughly 2–4× effective cost vs neoclouds — cost-optimal is not decision-optimal when compliance or capacity certainty is a hard constraint.
A significant fraction of current AI spend is subsidized — through provider credits, free tiers, academic programs, and below-cost pricing funded by venture capital — though the exact fraction is opaque. The effective cost of AI is materially higher than what most teams currently pay, and the subsidy distorts the build-vs-buy decision by making API inference appear cheaper than its true cost.
The verified price range, the quality adjustment, and the two largest non-routing levers.
Open this section on its ownThe inference-economics notebook has 17 references and 5 verified claims:
C1 (verified, 3 independent grade-A sources): The same API workload costs $0.018 to $2.00 across major LLMs — a 111x price range driven by the 21x blended-price spread in the frontier basket.
C2 (verified, 3 independent grade-A sources): Quality-adjusted cost diverges from token price by 47x across the frontier basket — far more than the 2x the original hypothesis predicted. DeepSeek V4-Flash at $2/index-point vs Claude Fable 5 at $94/index-point.
C3 (verified, 3 independent grade-A sources): The price-per-quality frontier is non-linear: models in the 40-50 Intelligence Index range offer 10-47x better cost-per-index-point than models in the 55-60 range. The marginal cost of capability accelerates above index 50.
C4 (verified, 1 grade-A source): Closed-weight flagships carry a 10% premium over open-weight equivalents in blended price ($4.47 vs $4.08/Mtok), but the premium concentrates at the frontier, not uniformly across tiers.
The decision as a sequence: measure effective cost, take the discounts, then route.
Open this section on its ownDon’t compare token prices — compare cost-per-quality-point for your workload. A model that is 10x cheaper per token but takes 3 retries to get a correct answer is more expensive than the one it replaced. The non-linear frontier means the biggest cost leverage is not picking the cheapest model, but picking the cheapest model that clears your quality bar — and that bar is workload-specific. Two caveats now carry the same weight as the rule itself: posted-price ranges are upper bounds (token-efficiency varies 2.65x+ by model), and the quality index behind any cost-per-quality ranking is an editorially- weighted composite — use it to find the non-linear region, not to rank models to a decimal.
Whether to buy routing or build it, given how thin the fee and the moat both are.
Open this section on its ownThe middleware that performs the routing is not where the money is, and may not be where it stays. OpenRouter — the category leader, processing over $100M in annualized inference spend and more than a quadrillion tokens/year by mid-2026 — charges a flat 5.5% on credit purchases with token prices passed through at provider list rates. Its fee is roughly one-tenth the size of the savings its own category claims to deliver. The layer is also structurally commoditizing: open-source gateways (LiteLLM) replicate the unified-API function free forever, 10+ commercial competitors ship the same feature set, and the standardized OpenAI-compatible interface makes switching cost nearly zero. Stripe’s $7.5B acquisition of OpenRouter (August 2026) reads as the category’s sustainability answer: routing survives as payments infrastructure, not as a standalone margin business.
Who absorbs the cost of model fit? Not the router. The 5.5% fee is the smallest line in the fit-cost stack; the real costs of matching models to workloads — evaluation effort to pick the threshold model, retry and quality-mismatch costs when the cheap model fails, re-integration each time a current model is deprecated — are absorbed by the developer. Routing-as-a- service removes the transport problem, not the fit problem.
What the index-dependence caveat does to the quality-adjusted figure you are about to quote.
Open this section on its ownHave you measured output-token volume per task on each candidate, not just the input price?
Does this workload reuse enough context for caching to beat the cache-write premium?
Would a 23-point index revision change which model you picked?
A routing and caching decision calculator. Specified in the audience ledger, not built — public tooling is board-gated, not blocked on research.
Routes are a reading order, not a substitute for the report. The full synthesis carries the update log and the complete evidence boundary.