Proof standard
Primary program terms and a cost-side margin anchor, not a vendor estimate.
Audience route
You are pricing a commitment whose posted price is not its cost.
The decision
Proof standard
Primary program terms and a cost-side margin anchor, not a vendor estimate.
Working output
A subsidy-expiry exposure list, one line per provider commitment.
The sections below are the report's own material, composed here. Nothing is rewritten for this audience; only the order and the framing are.
The subsidy finding, stated with the part of it that is still unmeasured.
Open this section on its ownThe sticker price of a token is not the cost of an answer. In August 2026, the same standard API workload (100K input + 20K output tokens) costs anywhere from $0.018 on Gemini Flash-Lite to $2.00 on Claude Fable 5 — a 111x price range across comparable models. Two corrections cut the other way: measured token-efficiency variance means the same input produces 2.65x+ more output tokens on some models, and cheap-token models can cost more per successful task via verbosity and retry loops — so the posted range is an upper bound on effective dispersion, and it still understates the real spread: adjusted for quality using the Artificial Analysis Intelligence Index, the cost-per-quality- point spreads by 47x. Read that 47x as order-of-magnitude evidence of non-linear pricing, not a precise ratio: the index is an editorially-weighted nine-benchmark composite (text/English-centric; agentic-weighted since v4.1), and version overhauls move scores by ~23 points — as large as the frontier spread itself. The marginal cost of capability accelerates sharply above Intelligence Index 50.
Running your own GPUs is a crossover problem, not a preference. Specialist GPU clouds (“neoclouds” like RunPod, Lambda, Together AI) charge 50-70% less than hyperscalers for the same H100: a median of $3.99/GPU-hour vs $7.89. The TCO model — now corrected for procurement scope — shows the answer splits by who’s buying: a lean operator adding GPUs to existing infrastructure (~$30K/GPU street price) breaks even against hyperscaler on-demand at ~21% utilization and beats median neocloud pricing above ~37%; an enterprise node-loaded buyer (~$94K/GPU all-in per io.net’s measured 3-year figure) breaks even vs hyperscaler on-demand only at ~57%, essentially never beats median neocloud pricing, and never beats cheap neoclouds or spot. Reserved commitments beat node-loaded on-prem at every utilization. Utilization assumption AND procurement scope together dominate the decision.
The control premium is real, and buyers pay it knowingly. Independent cost itemization puts true self-hosting at 3–5× the pure GPU price — matching the node-loaded math. For a compliance-driven subset, the premium is not optional: sovereign cloud IaaS spending is forecast at $80B in 2026 (+35.6%), Italy’s cloud market grew 20% YoY explicitly on sovereign-data demand, and the EU’s Technological Sovereignty Package (June 2026) extends data-governance obligations to AI providers. The other evidenced motives: guaranteed capacity with no rate limits or peak contention, latency for on-network workloads, vendor independence as a hedge against the list-price and subsidy moves forecast above, and very-high sustained utilization. In practice the control premium runs roughly 2–4× effective cost vs neoclouds — cost-optimal is not decision-optimal when compliance or capacity certainty is a hard constraint.
A significant fraction of current AI spend is subsidized — through provider credits, free tiers, academic programs, and below-cost pricing funded by venture capital — though the exact fraction is opaque. The effective cost of AI is materially higher than what most teams currently pay, and the subsidy distorts the build-vs-buy decision by making API inference appear cheaper than its true cost.
Orientation: plain-language definitions of the vocabulary this route's decisions use.
Open this section on its ownDecision-makers shouldn’t need an ML background to use this evidence. Plain-language definitions of the terms doing the heaviest work:
What to assume when the credit balance runs out.
Open this section on its ownInstrument cost per successful task, not cost per token. The 111x price range across models means a routing decision (cheap model for easy queries, expensive model for hard ones) is the single largest cost lever available — but the routing infrastructure itself has a cost that must be accounted for, and the savings are quality-conditioned: a measured 8-week pilot realized 58% cost reduction at a 91% response-acceptance rate, so the acceptance threshold you set — and the residual quality cost it implies — lands on you, not the router.
The dated subsidy-contraction bet and its probability.
Open this section on its ownThe forecast notebook has now asserted its dated binary bets (evidence cut 2026-08-20, resolution by August 2028):
| Bet | Claim | P | Confidence |
|---|---|---|---|
| FC1 | Model efficiency (not hardware) drives >50% of further price decline | 0.70 | medium |
| FC2 | Hyperscaler custom silicon reaches 25%+ of inference workload by mid-2028 | 0.45 | low-medium |
| FC3 | Jevons paradox holds — a 50% unit-cost cut raises volume more than 50% | 0.65 | medium |
| FC4 | Subsidies contract 50%+, raising effective cost 1.5–3x for subsidized users | 0.50 | low-medium |
| FC5 | Open-weight models reach quality parity on most workloads within 24 months | 0.55 | medium-low |
| FC6 | Edge becomes cost-advantageous for a materially larger workload set | 0.60 | medium-low |
The structural read: the 10x/year compression era is ending (fixed-quality decline is decelerating toward 1.5–5x/year and bifurcating — commodity approaching free, frontier reasoning moving up in price), so planning should treat unit cost as a shrinking but non-zero line item while total spend likely still rises (FC3). The least evidenced bets — subsidy contraction and demand elasticity — are the ones that would move budgets most.
For the shortest defensible account: three findings are durable on current evidence — the posted-price range is enormous and effective dispersion narrows only under task-level accounting (C1/C2, five independent sources each); owning GPUs is a procurement-scope decision with quantified crossover bands, not a preference; and the routing layer is thin-fee and commoditizing, with consolidation already begun. Three positions are bets, not facts — that unit- cost decline continues at 1.5–5x/year (FC1), that demand absorbs the savings (FC3), and that subsidies contract (FC4). The discipline for a reader: build on the durable findings, monitor the bets by their named triggers, and treat any plan that requires all six bets to resolve favorably as a plan with no margin.
Exactly which half of the subsidy claim is anchored and which is structural.
Open this section on its ownDoes any current unit economic depend on a credit balance or a free tier?
What is the switching cost if the zero-price channel caps change?
Is the lock-in priced, or only the rate?
A subsidy-expiry exposure checklist, with the zero-price channel caps rechecked quarterly.
Routes are a reading order, not a substitute for the report. The full synthesis carries the update log and the complete evidence boundary.