The real cost of AI: August 2026
Status: initial research snapshot. Evidence cut: 2026-08-20. Recheck by: 2026-11-20, or after a material model release, pricing change, or hardware generation shift.
This summary draws on the inference-economics notebook (17 captured references, 5 verified claims), the GPU ownership notebook (17 references including grade-A captures and a compiled 8-source pricing file, 4 verified claims plus the TCO crossover model), and the model-selection notebook (21 references, 6 verified claims incl. router economics and index-sensitivity findings). Counts as of 2026-08-21 (evening).
The short version
The sticker price of a token is not the cost of an answer. In August 2026, the same standard API workload (100K input + 20K output tokens) costs anywhere from $0.018 on Gemini Flash-Lite to $2.00 on Claude Fable 5 — a 111x price range across comparable models. Two corrections cut the other way: measured token-efficiency variance means the same input produces 2.65x+ more output tokens on some models, and cheap-token models can cost more per successful task via verbosity and retry loops — so the posted range is an upper bound on effective dispersion, and it still understates the real spread: adjusted for quality using the Artificial Analysis Intelligence Index, the cost-per-quality- point spreads by 47x. Read that 47x as order-of-magnitude evidence of non-linear pricing, not a precise ratio: the index is an editorially-weighted nine-benchmark composite (text/English-centric; agentic-weighted since v4.1), and version overhauls move scores by ~23 points — as large as the frontier spread itself. The marginal cost of capability accelerates sharply above Intelligence Index 50.
Running your own GPUs is a crossover problem, not a preference. Specialist GPU clouds (“neoclouds” like RunPod, Lambda, Together AI) charge 50-70% less than hyperscalers for the same H100: a median of $3.99/GPU-hour vs $7.89. The TCO model — now corrected for procurement scope — shows the answer splits by who’s buying: a lean operator adding GPUs to existing infrastructure (~$30K/GPU street price) breaks even against hyperscaler on-demand at ~21% utilization and beats median neocloud pricing above ~37%; an enterprise node-loaded buyer (~$94K/GPU all-in per io.net’s measured 3-year figure) breaks even vs hyperscaler on-demand only at ~57%, essentially never beats median neocloud pricing, and never beats cheap neoclouds or spot. Reserved commitments beat node-loaded on-prem at every utilization. Utilization assumption AND procurement scope together dominate the decision.
The control premium is real, and buyers pay it knowingly. Independent cost itemization puts true self-hosting at 3–5× the pure GPU price — matching the node-loaded math. For a compliance-driven subset, the premium is not optional: sovereign cloud IaaS spending is forecast at $80B in 2026 (+35.6%), Italy’s cloud market grew 20% YoY explicitly on sovereign-data demand, and the EU’s Technological Sovereignty Package (June 2026) extends data-governance obligations to AI providers. The other evidenced motives: guaranteed capacity with no rate limits or peak contention, latency for on-network workloads, vendor independence as a hedge against the list-price and subsidy moves forecast above, and very-high sustained utilization. In practice the control premium runs roughly 2–4× effective cost vs neoclouds — cost-optimal is not decision-optimal when compliance or capacity certainty is a hard constraint.
A significant fraction of current AI spend is subsidized — through provider credits, free tiers, academic programs, and below-cost pricing funded by venture capital — though the exact fraction is opaque. The effective cost of AI is materially higher than what most teams currently pay, and the subsidy distorts the build-vs-buy decision by making API inference appear cheaper than its true cost.
What the evidence establishes
Inference economics — five verified claims
The inference-economics notebook has 17 references and 5 verified claims:
-
C1 (verified, 3 independent grade-A sources): The same API workload costs $0.018 to $2.00 across major LLMs — a 111x price range driven by the 21x blended-price spread in the frontier basket.
-
C2 (verified, 3 independent grade-A sources): Quality-adjusted cost diverges from token price by 47x across the frontier basket — far more than the 2x the original hypothesis predicted. DeepSeek V4-Flash at $2/index-point vs Claude Fable 5 at $94/index-point.
-
C3 (verified, 3 independent grade-A sources): The price-per-quality frontier is non-linear: models in the 40-50 Intelligence Index range offer 10-47x better cost-per-index-point than models in the 55-60 range. The marginal cost of capability accelerates above index 50.
-
C4 (verified, 1 grade-A source): Closed-weight flagships carry a 10% premium over open-weight equivalents in blended price ($4.47 vs $4.08/Mtok), but the premium concentrates at the frontier, not uniformly across tiers.
GPU ownership — three verified claims and a crossover model
The GPU notebook has 7 references (a compiled 8-source evidence file grade C + direct captures including A6 grade A, A7 grade A) and 3 verified claims, plus the TCO crossover model (analysis/tco-crossover-2026-08.md):
-
C1 (verified, 3 independent sources): Specialist GPU clouds (“neoclouds” like RunPod, Lambda, Together AI) charge 50-70% less than hyperscalers for the same H100 GPU-hour: median $3.99/GPU-hr vs $7.89/GPU-hr (+98%). The ~2x hyperscaler premium is consistent across H100, A100, H200, and B200.
-
C2 (verified, 4 independent sources): On-premises GPU clusters break even with hyperscaler on-demand at roughly 50-83% sustained utilization, but rarely win against specialist GPU clouds at any utilization once fully loaded.
What this means for decisions
If you build with AI APIs
Don’t compare token prices — compare cost-per-quality-point for your workload. A model that is 10x cheaper per token but takes 3 retries to get a correct answer is more expensive than the one it replaced. The non-linear frontier means the biggest cost leverage is not picking the cheapest model, but picking the cheapest model that clears your quality bar — and that bar is workload-specific. Two caveats now carry the same weight as the rule itself: posted-price ranges are upper bounds (token-efficiency varies 2.65x+ by model), and the quality index behind any cost-per-quality ranking is an editorially- weighted composite — use it to find the non-linear region, not to rank models to a decimal.
If you run ML infrastructure
The GPU market is split: neoclouds at $3-4/GPU-hour, hyperscalers at $7-12. If you’re paying hyperscaler on-demand for sustained inference, you are likely overpaying by 2x. Whether owning beats renting now has a two-part answer. Procurement scope first: a lean team adding GPUs to existing infrastructure (~$30K/GPU) breaks even vs hyperscaler on-demand at ~21% utilization and beats median neocloud pricing above ~37% — but an enterprise node-loaded buyer (~$94K/GPU all-in) needs ~57% just against hyperscaler on-demand, essentially never beats median neoclouds, and never beats cheap neoclouds or spot. Utilization second: above ~70% sustained, on-prem wins on raw cost in the lean regime. And if you’re buying for sovereignty, compliance, or guaranteed capacity, you’re paying a 2–4x control premium on purpose — price it as insurance, not as infrastructure.
If you budget AI spend
Instrument cost per successful task, not cost per token. The 111x price range across models means a routing decision (cheap model for easy queries, expensive model for hard ones) is the single largest cost lever available — but the routing infrastructure itself has a cost that must be accounted for, and the savings are quality-conditioned: a measured 8-week pilot realized 58% cost reduction at a 91% response-acceptance rate, so the acceptance threshold you set — and the residual quality cost it implies — lands on you, not the router.
The 24-month horizon — six dated bets
The forecast notebook has now asserted its dated binary bets (evidence cut 2026-08-20, resolution by August 2028):
| Bet | Claim | P | Confidence |
|---|---|---|---|
| FC1 | Model efficiency (not hardware) drives >50% of further price decline | 0.70 | medium |
| FC2 | Hyperscaler custom silicon reaches 25%+ of inference workload by mid-2028 | 0.45 | low-medium |
| FC3 | Jevons paradox holds — a 50% unit-cost cut raises volume more than 50% | 0.65 | medium |
| FC4 | Subsidies contract 50%+, raising effective cost 1.5–3x for subsidized users | 0.50 | low-medium |
| FC5 | Open-weight models reach quality parity on most workloads within 24 months | 0.55 | medium-low |
| FC6 | Edge becomes cost-advantageous for a materially larger workload set | 0.60 | medium-low |
The structural read: the 10x/year compression era is ending (fixed-quality decline is decelerating toward 1.5–5x/year and bifurcating — commodity approaching free, frontier reasoning moving up in price), so planning should treat unit cost as a shrinking but non-zero line item while total spend likely still rises (FC3). The least evidenced bets — subsidy contraction and demand elasticity — are the ones that would move budgets most.
What’s durable vs what’s a bet
For the shortest defensible account: three findings are durable on current evidence — the posted-price range is enormous and effective dispersion narrows only under task-level accounting (C1/C2, five independent sources each); owning GPUs is a procurement-scope decision with quantified crossover bands, not a preference; and the routing layer is thin-fee and commoditizing, with consolidation already begun. Three positions are bets, not facts — that unit- cost decline continues at 1.5–5x/year (FC1), that demand absorbs the savings (FC3), and that subsidies contract (FC4). The discipline for a reader: build on the durable findings, monitor the bets by their named triggers, and treat any plan that requires all six bets to resolve favorably as a plan with no margin.
The routing layer itself — thin fees, thin moats
The middleware that performs the routing is not where the money is, and may not be where it stays. OpenRouter — the category leader, processing over $100M in annualized inference spend and more than a quadrillion tokens/year by mid-2026 — charges a flat 5.5% on credit purchases with token prices passed through at provider list rates. Its fee is roughly one-tenth the size of the savings its own category claims to deliver. The layer is also structurally commoditizing: open-source gateways (LiteLLM) replicate the unified-API function free forever, 10+ commercial competitors ship the same feature set, and the standardized OpenAI-compatible interface makes switching cost nearly zero. Stripe’s $7.5B acquisition of OpenRouter (August 2026) reads as the category’s sustainability answer: routing survives as payments infrastructure, not as a standalone margin business.
Who absorbs the cost of model fit? Not the router. The 5.5% fee is the smallest line in the fit-cost stack; the real costs of matching models to workloads — evaluation effort to pick the threshold model, retry and quality-mismatch costs when the cheap model fails, re-integration each time a current model is deprecated — are absorbed by the developer. Routing-as-a- service removes the transport problem, not the fit problem.
Exit checklist — how to keep switching cost near zero. Since the layer’s standardization is what makes switching cheap, portability is a property you keep or lose by construction. Six checks, all cheap to maintain from day one:
- Speak only the OpenAI-compatible schema at your client boundary — no provider-native SDK types leaking past it.
- Keep the eval harness router-independent: acceptance thresholds measured against model outputs directly, so re-routing never invalidates your bar.
- Export usage/cost logs to storage you own — billing data is the lever in any renegotiation.
- Avoid provider-exclusive parameters (fine-grained logit control, native tool dialects) unless the dependency is worth a migration.
- Re-run the quality threshold quarterly against the current frontier basket — the non-linear frontier means last quarter’s “cheap model that clears” may no longer be on it.
- Price the exit annually: hours to swap gateways × loaded rate. If that number grows, the middleware has become infrastructure — renegotiate or migrate before it becomes both.
If you price an AI-powered product
Cost-per-user is three multiplications, and the subsidy question dominates all of them. First, tokens per active user: your workload’s input/output profile times usage frequency — measure it, don’t estimate from demos, because token- efficiency varies 2.65x+ between models on the same input. Second, the rate: the frontier’s 111x posted range means model choice sets the rate, and caching (50–90% off repeated context) plus batch (flat 50%) can cut it again — with the cache-write caveat that one-shot workloads may not benefit. Third, the subsidy exposure: if any of your cost base rides zero-price channels (:free variants) or startup credits, your unit economics are temporary by construction — price the product at list, and treat today’s subsidy as margin, not as your cost basis.
The failure mode this session’s evidence most warns about: a product whose gross margin works only while someone else funds the inference. When the subsidy contracts (FC4), the teams harmed are precisely those who priced against subsidized rates without a list-price floor in their model.
If you’re evaluating fine-tuning vs retrieval
Three verified findings from the fine-tuning notebook, all pointing the same direction. PEFT (LoRA/QLoRA) makes the compute almost free — a 70B QLoRA run costs tens of dollars on a single H100 vs 8×H100 hours for full fine-tuning. Compute is nonetheless only 2–3% of a real programme: dataset curation, evaluation, and MLOps dominate, and two conditions widen that gap — fine-tuning causes unpredictable safety drift even on clean data (budget recurring safety evals per refit), and every base-model deprecation forces a refit, making compute a subscription rather than a purchase. On the build decision itself: RAG and fine-tuning are substitutes for knowledge-heavy tasks (RAG is updateable and cheaper to iterate) but complements for behavior-heavy tasks — fine-tune the behavior, retrieve the knowledge.
Setting your 2027 AI budget — the six-decision sequence
If you are writing a 2027 AI budget between now and January, the evidence in this report maps onto six sequential decisions. Each names the job, what the verified evidence says, and what would change the answer.
1. Set the envelope: assume unit cost falls, total spend rises. Fixed- quality inference is still getting cheaper (1.5–5x/year and decelerating), but the Jevons bet (FC3, p=0.65) says volume grows faster than price falls. Budget the unit line shrinking and the volume line growing; a budget that assumes both flat will be spent differently than written by Q2.
2. Choose your commitment structure: treat subsidies as expiring. The cost-side anchor is verified — frontier list prices run below sustainable infrastructure return (OpenAI’s gross margin fell to 33% against its own 46% forecast). The subsidy-contraction bet (FC4, p=0.50) says effective cost for subsidized teams rises 1.5–3x over the window. Practical rule: build your base budget at list price with no credits, then treat credits as a discount to be won, not a floor to stand on. Audit which of your workloads sit on zero-price channels (:free variants, launch access) — that channel bites first if contraction happens.
3. Decide capacity ownership by procurement scope, not by preference. The two-regime answer: lean buyers adding GPUs to existing infrastructure break even vs hyperscaler on-demand at ~21% utilization and beat median neoclouds above ~37%; enterprise node-loaded buyers (~$94K/GPU all-in) need ~57% and effectively never beat neoclouds or spot. If you’re buying anyway for sovereignty, compliance, or guaranteed capacity, price it as insurance — roughly 2–4x effective cost — because regulatory sovereignty demand is measurably expanding ($80B sovereign IaaS forecast for 2026, EU Sovereignty Package in force).
4. Structure the routing/middleware stack for exit. The layer charges a thin flat fee (~5.5%), delivers savings an order of magnitude larger, and is decommoditizing under your feet — open-source gateways replicate it free, switching costs are near zero, and consolidation has begun (Stripe/OpenRouter). Use routers for transport; keep model choice, evaluation thresholds, and prompt architecture portable. The fit problem — knowing which model clears your bar, and paying when it doesn’t — stays yours no matter what you route through.
5. Budget evaluation explicitly. The routing-savings evidence is quality- conditioned (58% realized at 91% acceptance in the measured pilot), and the quality index behind any comparison is editorially weighted. Your acceptance threshold is a business decision that needs its own eval budget — the line item most teams forget, and the one that determines whether the routing savings are real.
6. Write the reopen conditions into the budget itself. Every number above is evidence-dated 2026-08-20/21. Name what reopens each line: a material repricing or model release, subsidy program changes, :free channel cap moves, or the quarterly capex map. A budget with reopen conditions survives contact with the market; one without them gets defended after it’s wrong.
Worked illustration — a 20-engineer product team
Illustrative arithmetic from the custody rates above (not a claim about any team; substitute your own volumes). Suppose the team spent ~$96K on inference in 2026 at blended list rates, has no compliance constraints, and is writing its 2027 budget:
- Envelope: unit costs fall 2–4x for fixed quality (FC1/FC2 trajectory); volume grows faster (FC3). Line item: $60–80K at list for more usage than 2026 — not $96K for the same usage.
- Commitments: no multi-year lock-ins; credits booked as upside, not base. Exposure if :free channels or startup credits vanish: priced at zero.
- Capacity ownership: at this scale, procurement scope is enterprise-grade if bought new (~$94K/GPU all-in) — ownership loses to neoclouds at every utilization below ~99%. Answer: none. Revisit only if sustained utilization projection exceeds ~70% and hardware is already sunk.
- Routing stack: one flat-fee aggregator or a self-hosted gateway behind an abstraction boundary; model choice re-rankable monthly. Budget line: ~$0 for transport, ~$6–10K for evaluation (the acceptance-threshold work).
- Evaluation budget: the acceptance threshold from the measured pilot is the difference between 58% savings and silent quality loss. Fund it as a first-class line, not spare time.
- Reopen conditions written in: repricing >20% by any incumbent; FC4 evidence (credit-program contraction); Rubin shipment data moving GPU rental prices >25%.
The pattern generalizes: at small-to-mid scale the budget is mostly API line items plus an eval budget; ownership enters only through compliance or very high utilization, and middleware should never be a lock-in.
Worked illustration — the compliance-constrained variant
Same team size, but the workload processes regulated personal data that cannot leave a defined jurisdiction, and projected utilization is high (~70% sustained) because the inference serves a core product loop. Now the decision tree changes shape:
- The API-first answer is not disqualified by price — it’s disqualified by constraint. If no provider region satisfies the residency requirement, the comparison starts between sovereign-hosted and owned.
- Sovereign/regional hosting: expect regional-provider rates at or above median neocloud pricing ($4–6/GPU-hour effective in custody examples) plus residency premiums. At 70% utilization that’s roughly $1,900–2,900/GPU-month.
- Owned, node-loaded: ~$2,900/GPU-month fully loaded at 100% utilization (io.net basis), scaling down modestly with idle power. At ~70% utilization the two lanes reach parity — which is exactly why sovereignty-era buyers are choosing ownership at utilizations the pure-cost model calls marginal.
- What tips it: guaranteed capacity (no contention at peak), auditability of every token path, and freedom from list-price moves (FC4’s upside case makes this worth more, not less).
- What to budget anyway: the safety-drift and refit conditions apply twice over — owned models still deprecate, and compliance regimes require documented evaluation per refit.
The general rule for constrained buyers: the control premium is not a cost overrun, it’s the price of the constraint — and at high sustained utilization it approaches zero. The teams overpaying are the ones who buy the premium without having the constraint.
Context — the vocabulary this report assumes
Decision-makers shouldn’t need an ML background to use this evidence. Plain-language definitions of the terms doing the heaviest work:
- Token — the unit of text models read and write (~¾ of a word). All pricing starts here, which is why token counts, not just rates, matter.
- PEFT / LoRA / QLoRA — techniques that adapt a large model by training a tiny fraction of it. Makes fine-tuning compute cheap enough to run on one GPU — but doesn’t shrink the surrounding programme costs (data, evaluation, refits).
- Neocloud — a specialist GPU cloud (RunPod, Lambda, Together AI) renting raw GPU-hours at roughly half of hyperscaler prices, typically without the enterprise contracting layer.
- On-demand / reserved / spot — three rental tiers: pay-per-hour at list; 30–62% off for committing 1–3 years; and 60–91% off with interruption risk (the machine can be taken back mid-job).
- Prompt caching — paying once to process a repeated context (system prompts, documents), then 50–90% less to re-read it. Cache writes cost 1.25–2× base, so it only pays when context repeats.
- Batch API — half price on all tokens in exchange for results within 24 hours instead of seconds.
- Blended price — input and output rates combined into one number using a fixed ratio, so different models can be compared on one axis.
- Cost per successful task — what this report argues you should actually measure: total tokens (including retries and verbosity) divided by tasks completed acceptably. The only unit that reflects quality differences.
- Artificial Analysis Intelligence Index — an independent composite score (nine benchmarks, editorially weighted, text/English-centric) used here as the quality axis. Useful for locating the non-linear region; not precise enough to rank models to a decimal.
- Jevons paradox — the pattern where cheaper units increase total consumption enough that total spend rises. Bet FC3 asserts it applies here.
- Control premium — the 2–4× effective cost sovereign/compliance buyers knowingly pay for owning infrastructure rather than renting it.
- Procurement scope — whether “owning” means buying GPUs into existing infrastructure (~$30K/GPU) or procuring complete nodes with networking and financing (~$94K/GPU). The single assumption that swings every ownership crossover band.
- Utilization — the share of an owned GPU’s time doing paid work. The load-bearing variable in every own-vs-rent comparison.
- :free variants — aggregator-hosted model endpoints priced at literally $0/token with request caps. Real usage flows through them; they are also the first thing subsidy contraction would remove.
Evidence boundaries
- All prices are evidence of posted rates on 2026-08-20, not durable claims. Provider pricing changes weekly.
- The quality-adjusted cost analysis uses the Artificial Analysis Intelligence Index, which is a composite benchmark score — it does not predict performance on a specific workload. The index is also an editorially-weighted nine-benchmark composite (text/English-centric, agentic-shifted since v4.1); version overhauls have moved scores by ~23 points, so cost-per-index-point figures are order-of-magnitude evidence, not precise rankings.
- The GPU pricing data is compiled from independent sources including two grade-A direct captures (getdeploying.com, spendark.com) and a compiled evidence file; both GPU claims are verified.
- The subsidy analysis is structural (the subsidy exists, it distorts decisions) and now has a cost-side quantitative anchor: frontier-lab list pricing runs below sustainable infrastructure return (OpenAI gross margin 33% in 2025 vs its own 46% forecast, inference costs ~$8.4B and 4x YoY; Anthropic 40%) — posted API prices remain effectively investor-subsidized. The user-side subsidized fraction is still unmeasured.
- The forecast bets above are dated, binary, probability-bearing judgments over a short record — not calibrated forecasts. FC4 and FC6 have named missing baselines; their probabilities are structured priors.
- Recheck by: 2026-11-20, or after a material model release, pricing change, or hardware generation shift.
Update log
-
2026-08-21 (14): Added the missing fine-tuning decision layer (researcher route): PEFT economics, compute-as-subscription conditions (safety-drift evals per refit; base-model deprecation refits), and the RAG-substitutes/complements split — from the finetuning notebook’s three verified claims, previously absent from the report.
-
2026-08-21 (13): Decision-guidance sections updated to match current evidence: infrastructure advice rewritten around the two-regime TCO answer (replacing the superseded 50–83% single-band figure) with control-premium guidance; API-builder advice carries the token-efficiency and index-weighting caveats.
-
2026-08-21 (12): Adversarial pass on the 47x cost-per-quality spread: verified C6 (model-selection notebook) — the AA Intelligence Index is an editorially-weighted, text/English-centric composite whose v4.1 reweighting moved scores by ~23 points (comparable to the frontier spread itself). The 47x is now presented as order-of-magnitude evidence of non-linear pricing, not a precise ratio.
-
2026-08-21 (11): Verified C4 (GPU notebook): regulatory sovereignty drives on-prem/sovereign AI infrastructure independent of cost — $80B 2026 sovereign IaaS forecast (+35.6%), Italy +20% YoY on sovereign-data demand, EU Sovereignty Package (June 2026). Compliance buyers pay the control premium as a requirement. Also fixed a custody violation (Italy figure cited from a search snippet before capture).
-
2026-08-21 (10): Control-premium analysis added to the GPU section: independent itemization (Vensas, grade A) confirms self-hosting at 3–5× pure GPU price, independently corroborating the io.net node-loaded correction; five evidenced motives for paying it (sovereignty/compliance, guaranteed capacity, latency, vendor independence, sustained high utilization); premium quantified at roughly 2–4× effective cost vs neoclouds.
-
2026-08-21 (9): Adversarial pass on the TCO model found a material flaw: our bands used GPU-street-price capex ($30K/GPU), omitting node integration and financing. With io.net’s measured enterprise figure ($93,750/GPU over 3yr), on-prem crossovers move dramatically (vs neocloud median: ~37% → ~99%, i.e., effectively never; vs spot: never) — reconciling previously contradictory custody claims as population differences (lean GPU-only buyers vs enterprise node-loaded buyers). Model runs both scenarios; report section rewritten.
-
2026-08-21 (8): Adversarial pass on the 111x headline: no source disputes the posted-price spread, but measured token-efficiency variance (TensorZero, grade A: 2.65x+ same-input token-count difference) and retry/verbosity costs mean the posted range is an upper bound on effective dispersion. Headline and claim prose updated; C1 corroboration now 5.
-
2026-08-21 (7): Adversarial pass on the routing-savings claim: hunted counterevidence, found quality-accounted confirmation instead (RouteNLP 8-week pilot: 58% cost reduction at 91% acceptance; corroboration now 5). Claim prose and report passage updated with the acceptance-rate condition.
-
2026-08-21 (6): New reader-commissioned angle: the routing/middleware layer itself. Verified C4 (thin flat take-rate — OpenRouter 5.5%, >$100M annualized spend routed) and C5 (structural commoditization — free self-host gateways, near-zero switching costs, Stripe’s $7.5B acquisition as the sustainability answer). Added “who absorbs model-fit cost” analysis: developers, not routers.
-
2026-08-21 (5): Custody audit of this report: notebook inventory counts refreshed to current custody (15 refs / 5 claims inference-economics; 7 refs / 3 claims + TCO model GPU; 8 refs / 3 claims model-selection); spot-risk scenario computed in the TCO model (spot-vs-on-prem crossover ~67% → ~47% at a 20% interruption premium — replaces the earlier estimated figure); C5 cache-write caveat added.
-
2026-08-21 (4): Evidence-strength pass: all seven previously single-corroboration load-bearing claims brought to ≥2 independent sources — closed-weight premium (DeepInfra), caching/batch discounts (Amnic, pecollective), routing savings (leanlm, sirayatech), index spread (swfte composite), NVIDIA cadence (ModulEdge/Kuo + official PR), edge NPU price/perf (Hackster.io + official), startup credit amounts (AWS/GCP primary terms + pragma-code comparison). No conclusions changed; confidence strengthened.
-
2026-08-21 (3): Built the GPU TCO crossover model (gpu-003): crossover bands vs on-prem quantified per lane — hyperscaler on-demand ~21%, 1-yr reserved ~30%, 3-yr ~43%, neocloud median ~37%/low ~51%, spot ~67% utilization; workload views added (always-on → on-prem; business-hours → spot). Refined the earlier “on-prem rarely wins vs neoclouds” statement into a named-assumption band (amortization window and staffing move it across ~35–65%).
-
2026-08-21 (2): Captured the missing forecast baselines: verified C3 in the subsidies notebook (frontier list pricing below sustainable infrastructure return — OpenAI GM 33% vs 46% forecast, inference ~$8.4B 4x YoY; Anthropic 40%) and C4 in the edge notebook (current-gen NPU throughput at consumer prices: $249 Orin Nano Super ~35–54 tok/s on 1B-class models; AGX Thor 41.3 tok/s Llama 3.1 8B, 61 tok/s Qwen3 30B-A3B). Forecast bets FC4/FC6 upgraded from pure priors to informed bets; subsidy evidence boundary updated.
-
2026-08-21: Asserted the six dated forecast bets (FC1–FC6) from the ai-cost-forecast notebook; replaced the “no forecast yet” boundary with the bet table and its uncertainty statement. No price claims changed.
-
2026-08-20: Initial research snapshot. First public report. Drew on inference-economics (then 7 sources, 4 verified claims) and GPU-ownership (then 4 sources, 2 verified claims) notebooks. GPU sources re-captured via –via direct after default lane failed on cistern miss; record-only sources superseded and passed.