Proof standard
A payback volume, plus the retraining cadence that resets it.
Audience route
You are deciding whether a narrower model earns the work it takes to make one.
The decision
Proof standard
A payback volume, plus the retraining cadence that resets it.
Working output
A break-even task volume for the narrow model against the general one.
The sections below are the report's own material, composed here. Nothing is rewritten for this audience; only the order and the framing are.
The cost frame this decision sits inside.
Open this section on its ownThe sticker price of a token is not the cost of an answer. In August 2026, the same standard API workload (100K input + 20K output tokens) costs anywhere from $0.018 on Gemini Flash-Lite to $2.00 on Claude Fable 5 — a 111x price range across comparable models. Two corrections cut the other way: measured token-efficiency variance means the same input produces 2.65x+ more output tokens on some models, and cheap-token models can cost more per successful task via verbosity and retry loops — so the posted range is an upper bound on effective dispersion, and it still understates the real spread: adjusted for quality using the Artificial Analysis Intelligence Index, the cost-per-quality- point spreads by 47x. Read that 47x as order-of-magnitude evidence of non-linear pricing, not a precise ratio: the index is an editorially-weighted nine-benchmark composite (text/English-centric; agentic-weighted since v4.1), and version overhauls move scores by ~23 points — as large as the frontier spread itself. The marginal cost of capability accelerates sharply above Intelligence Index 50.
Running your own GPUs is a crossover problem, not a preference. Specialist GPU clouds (“neoclouds” like RunPod, Lambda, Together AI) charge 50-70% less than hyperscalers for the same H100: a median of $3.99/GPU-hour vs $7.89. The TCO model — now corrected for procurement scope — shows the answer splits by who’s buying: a lean operator adding GPUs to existing infrastructure (~$30K/GPU street price) breaks even against hyperscaler on-demand at ~21% utilization and beats median neocloud pricing above ~37%; an enterprise node-loaded buyer (~$94K/GPU all-in per io.net’s measured 3-year figure) breaks even vs hyperscaler on-demand only at ~57%, essentially never beats median neocloud pricing, and never beats cheap neoclouds or spot. Reserved commitments beat node-loaded on-prem at every utilization. Utilization assumption AND procurement scope together dominate the decision.
The control premium is real, and buyers pay it knowingly. Independent cost itemization puts true self-hosting at 3–5× the pure GPU price — matching the node-loaded math. For a compliance-driven subset, the premium is not optional: sovereign cloud IaaS spending is forecast at $80B in 2026 (+35.6%), Italy’s cloud market grew 20% YoY explicitly on sovereign-data demand, and the EU’s Technological Sovereignty Package (June 2026) extends data-governance obligations to AI providers. The other evidenced motives: guaranteed capacity with no rate limits or peak contention, latency for on-network workloads, vendor independence as a hedge against the list-price and subsidy moves forecast above, and very-high sustained utilization. In practice the control premium runs roughly 2–4× effective cost vs neoclouds — cost-optimal is not decision-optimal when compliance or capacity certainty is a hard constraint.
A significant fraction of current AI spend is subsidized — through provider credits, free tiers, academic programs, and below-cost pricing funded by venture capital — though the exact fraction is opaque. The effective cost of AI is materially higher than what most teams currently pay, and the subsidy distorts the build-vs-buy decision by making API inference appear cheaper than its true cost.
Orientation: plain-language definitions of the vocabulary this route's decisions use.
Open this section on its ownDecision-makers shouldn’t need an ML background to use this evidence. Plain-language definitions of the terms doing the heaviest work:
Where the compute cost stops and the real cost starts.
Open this section on its ownThree verified findings from the fine-tuning notebook, all pointing the same direction. PEFT (LoRA/QLoRA) makes the compute almost free — a 70B QLoRA run costs tens of dollars on a single H100 vs 8×H100 hours for full fine-tuning. Compute is nonetheless only 2–3% of a real programme: dataset curation, evaluation, and MLOps dominate, and two conditions widen that gap — fine-tuning causes unpredictable safety drift even on clean data (budget recurring safety evals per refit), and every base-model deprecation forces a refit, making compute a subscription rather than a purchase. On the build decision itself: RAG and fine-tuning are substitutes for knowledge-heavy tasks (RAG is updateable and cheaper to iterate) but complements for behavior-heavy tasks — fine-tune the behavior, retrieve the knowledge.
The per-token baseline the narrow model has to beat in production, not in training.
Open this section on its ownThe inference-economics notebook has 17 references and 5 verified claims:
C1 (verified, 3 independent grade-A sources): The same API workload costs $0.018 to $2.00 across major LLMs — a 111x price range driven by the 21x blended-price spread in the frontier basket.
C2 (verified, 3 independent grade-A sources): Quality-adjusted cost diverges from token price by 47x across the frontier basket — far more than the 2x the original hypothesis predicted. DeepSeek V4-Flash at $2/index-point vs Claude Fable 5 at $94/index-point.
C3 (verified, 3 independent grade-A sources): The price-per-quality frontier is non-linear: models in the 40-50 Intelligence Index range offer 10-47x better cost-per-index-point than models in the 55-60 range. The marginal cost of capability accelerates above index 50.
C4 (verified, 1 grade-A source): Closed-weight flagships carry a 10% premium over open-weight equivalents in blended price ($4.47 vs $4.08/Mtok), but the premium concentrates at the frontier, not uniformly across tiers.
What the fine-tuning evidence does not settle.
Open this section on its ownHave you costed data preparation and evaluation, or only GPU hours?
Is the task knowledge-heavy or behavior-heavy — substitutes or complements?
How often does the underlying knowledge change, and what does that do to payback?
A payback calculator that takes retraining cadence as an input.
Routes are a reading order, not a substitute for the report. The full synthesis carries the update log and the complete evidence boundary.