Proof standard
A cost per user computed from measured token volume, with the subsidy assumption named.
Audience route
You are checking whether an AI feature's unit economics hold at scale.
The decision
Proof standard
A cost per user computed from measured token volume, with the subsidy assumption named.
Working output
A cost-per-user model carrying an explicit subsidy-withdrawal scenario.
The sections below are the report's own material, composed here. Nothing is rewritten for this audience; only the order and the framing are.
The two findings that most often break an AI feature's margin model.
Open this section on its ownThe sticker price of a token is not the cost of an answer. In August 2026, the same standard API workload (100K input + 20K output tokens) costs anywhere from $0.018 on Gemini Flash-Lite to $2.00 on Claude Fable 5 — a 111x price range across comparable models. Two corrections cut the other way: measured token-efficiency variance means the same input produces 2.65x+ more output tokens on some models, and cheap-token models can cost more per successful task via verbosity and retry loops — so the posted range is an upper bound on effective dispersion, and it still understates the real spread: adjusted for quality using the Artificial Analysis Intelligence Index, the cost-per-quality- point spreads by 47x. Read that 47x as order-of-magnitude evidence of non-linear pricing, not a precise ratio: the index is an editorially-weighted nine-benchmark composite (text/English-centric; agentic-weighted since v4.1), and version overhauls move scores by ~23 points — as large as the frontier spread itself. The marginal cost of capability accelerates sharply above Intelligence Index 50.
Running your own GPUs is a crossover problem, not a preference. Specialist GPU clouds (“neoclouds” like RunPod, Lambda, Together AI) charge 50-70% less than hyperscalers for the same H100: a median of $3.99/GPU-hour vs $7.89. The TCO model — now corrected for procurement scope — shows the answer splits by who’s buying: a lean operator adding GPUs to existing infrastructure (~$30K/GPU street price) breaks even against hyperscaler on-demand at ~21% utilization and beats median neocloud pricing above ~37%; an enterprise node-loaded buyer (~$94K/GPU all-in per io.net’s measured 3-year figure) breaks even vs hyperscaler on-demand only at ~57%, essentially never beats median neocloud pricing, and never beats cheap neoclouds or spot. Reserved commitments beat node-loaded on-prem at every utilization. Utilization assumption AND procurement scope together dominate the decision.
The control premium is real, and buyers pay it knowingly. Independent cost itemization puts true self-hosting at 3–5× the pure GPU price — matching the node-loaded math. For a compliance-driven subset, the premium is not optional: sovereign cloud IaaS spending is forecast at $80B in 2026 (+35.6%), Italy’s cloud market grew 20% YoY explicitly on sovereign-data demand, and the EU’s Technological Sovereignty Package (June 2026) extends data-governance obligations to AI providers. The other evidenced motives: guaranteed capacity with no rate limits or peak contention, latency for on-network workloads, vendor independence as a hedge against the list-price and subsidy moves forecast above, and very-high sustained utilization. In practice the control premium runs roughly 2–4× effective cost vs neoclouds — cost-optimal is not decision-optimal when compliance or capacity certainty is a hard constraint.
A significant fraction of current AI spend is subsidized — through provider credits, free tiers, academic programs, and below-cost pricing funded by venture capital — though the exact fraction is opaque. The effective cost of AI is materially higher than what most teams currently pay, and the subsidy distorts the build-vs-buy decision by making API inference appear cheaper than its true cost.
Orientation: plain-language definitions of the vocabulary this route's decisions use.
Open this section on its ownDecision-makers shouldn’t need an ML background to use this evidence. Plain-language definitions of the terms doing the heaviest work:
The levers available before you change the product.
Open this section on its ownDon’t compare token prices — compare cost-per-quality-point for your workload. A model that is 10x cheaper per token but takes 3 retries to get a correct answer is more expensive than the one it replaced. The non-linear frontier means the biggest cost leverage is not picking the cheapest model, but picking the cheapest model that clears your quality bar — and that bar is workload-specific. Two caveats now carry the same weight as the rule itself: posted-price ranges are upper bounds (token-efficiency varies 2.65x+ by model), and the quality index behind any cost-per-quality ranking is an editorially- weighted composite — use it to find the non-linear region, not to rank models to a decimal.
The envelope the margin model has to live inside.
Open this section on its ownInstrument cost per successful task, not cost per token. The 111x price range across models means a routing decision (cheap model for easy queries, expensive model for hard ones) is the single largest cost lever available — but the routing infrastructure itself has a cost that must be accounted for, and the savings are quality-conditioned: a measured 8-week pilot realized 58% cost reduction at a 91% response-acceptance rate, so the acceptance threshold you set — and the residual quality cost it implies — lands on you, not the router.
Which direction the inputs move over a funding cycle.
Open this section on its ownThe forecast notebook has now asserted its dated binary bets (evidence cut 2026-08-20, resolution by August 2028):
| Bet | Claim | P | Confidence |
|---|---|---|---|
| FC1 | Model efficiency (not hardware) drives >50% of further price decline | 0.70 | medium |
| FC2 | Hyperscaler custom silicon reaches 25%+ of inference workload by mid-2028 | 0.45 | low-medium |
| FC3 | Jevons paradox holds — a 50% unit-cost cut raises volume more than 50% | 0.65 | medium |
| FC4 | Subsidies contract 50%+, raising effective cost 1.5–3x for subsidized users | 0.50 | low-medium |
| FC5 | Open-weight models reach quality parity on most workloads within 24 months | 0.55 | medium-low |
| FC6 | Edge becomes cost-advantageous for a materially larger workload set | 0.60 | medium-low |
The structural read: the 10x/year compression era is ending (fixed-quality decline is decelerating toward 1.5–5x/year and bifurcating — commodity approaching free, frontier reasoning moving up in price), so planning should treat unit cost as a shrinking but non-zero line item while total spend likely still rises (FC3). The least evidenced bets — subsidy contraction and demand elasticity — are the ones that would move budgets most.
For the shortest defensible account: three findings are durable on current evidence — the posted-price range is enormous and effective dispersion narrows only under task-level accounting (C1/C2, five independent sources each); owning GPUs is a procurement-scope decision with quantified crossover bands, not a preference; and the routing layer is thin-fee and commoditizing, with consolidation already begun. Three positions are bets, not facts — that unit- cost decline continues at 1.5–5x/year (FC1), that demand absorbs the savings (FC3), and that subsidies contract (FC4). The discipline for a reader: build on the durable findings, monitor the bets by their named triggers, and treat any plan that requires all six bets to resolve favorably as a plan with no margin.
The six-decision sequence for writing a 2027 AI budget — envelope, commitments, capacity ownership, routing portability, eval budget, reopen conditions.
Open this section on its ownIf you are writing a 2027 AI budget between now and January, the evidence in this report maps onto six sequential decisions. Each names the job, what the verified evidence says, and what would change the answer.
1. Set the envelope: assume unit cost falls, total spend rises. Fixed- quality inference is still getting cheaper (1.5–5x/year and decelerating), but the Jevons bet (FC3, p=0.65) says volume grows faster than price falls. Budget the unit line shrinking and the volume line growing; a budget that assumes both flat will be spent differently than written by Q2.
2. Choose your commitment structure: treat subsidies as expiring. The cost-side anchor is verified — frontier list prices run below sustainable infrastructure return (OpenAI’s gross margin fell to 33% against its own 46% forecast). The subsidy-contraction bet (FC4, p=0.50) says effective cost for subsidized teams rises 1.5–3x over the window. Practical rule: build your base budget at list price with no credits, then treat credits as a discount to be won, not a floor to stand on. Audit which of your workloads sit on zero-price channels (:free variants, launch access) — that channel bites first if contraction happens.
3. Decide capacity ownership by procurement scope, not by preference. The two-regime answer: lean buyers adding GPUs to existing infrastructure break even vs hyperscaler on-demand at ~21% utilization and beat median neoclouds above ~37%; enterprise node-loaded buyers (~$94K/GPU all-in) need ~57% and effectively never beat neoclouds or spot. If you’re buying anyway for sovereignty, compliance, or guaranteed capacity, price it as insurance — roughly 2–4x effective cost — because regulatory sovereignty demand is measurably expanding ($80B sovereign IaaS forecast for 2026, EU Sovereignty Package in force).
4. Structure the routing/middleware stack for exit. The layer charges a thin flat fee (~5.5%), delivers savings an order of magnitude larger, and is decommoditizing under your feet — open-source gateways replicate it free, switching costs are near zero, and consolidation has begun (Stripe/OpenRouter). Use routers for transport; keep model choice, evaluation thresholds, and prompt architecture portable. The fit problem — knowing which model clears your bar, and paying when it doesn’t — stays yours no matter what you route through.
5. Budget evaluation explicitly. The routing-savings evidence is quality- conditioned (58% realized at 91% acceptance in the measured pilot), and the quality index behind any comparison is editorially weighted. Your acceptance threshold is a business decision that needs its own eval budget — the line item most teams forget, and the one that determines whether the routing savings are real.
6. Write the reopen conditions into the budget itself. Every number above is evidence-dated 2026-08-20/21. Name what reopens each line: a material repricing or model release, subsidy program changes, :free channel cap moves, or the quarterly capex map. A budget with reopen conditions survives contact with the market; one without them gets defended after it’s wrong.
Illustrative arithmetic from the custody rates above (not a claim about any team; substitute your own volumes). Suppose the team spent ~$96K on inference in 2026 at blended list rates, has no compliance constraints, and is writing its 2027 budget:
The pattern generalizes: at small-to-mid scale the budget is mostly API line items plus an eval budget; ownership enters only through compliance or very high utilization, and middleware should never be a lock-in.
Same team size, but the workload processes regulated personal data that cannot leave a defined jurisdiction, and projected utilization is high (~70% sustained) because the inference serves a core product loop. Now the decision tree changes shape:
The general rule for constrained buyers: the control premium is not a cost overrun, it’s the price of the constraint — and at high sustained utilization it approaches zero. The teams overpaying are the ones who buy the premium without having the constraint.
The cost-per-user arithmetic and the subsidy-expiry failure mode, for pricing AI features.
Open this section on its ownCost-per-user is three multiplications, and the subsidy question dominates all of them. First, tokens per active user: your workload’s input/output profile times usage frequency — measure it, don’t estimate from demos, because token- efficiency varies 2.65x+ between models on the same input. Second, the rate: the frontier’s 111x posted range means model choice sets the rate, and caching (50–90% off repeated context) plus batch (flat 50%) can cut it again — with the cache-write caveat that one-shot workloads may not benefit. Third, the subsidy exposure: if any of your cost base rides zero-price channels (:free variants) or startup credits, your unit economics are temporary by construction — price the product at list, and treat today’s subsidy as margin, not as your cost basis.
The failure mode this session’s evidence most warns about: a product whose gross margin works only while someone else funds the inference. When the subsidy contracts (FC4), the teams harmed are precisely those who priced against subsidized rates without a list-price floor in their model.
Is cost per user measured from real traffic or estimated from a rate card?
Does the margin hold if effective cost rises 1.5–3x?
Which retry and verbosity behavior is inside your cost model?
A worked cost-per-user case with the retry loop included.
Routes are a reading order, not a substitute for the report. The full synthesis carries the update log and the complete evidence boundary.