Audience route

Researchers and data scientists

You are deciding whether a narrower model earns the work it takes to make one.

Whether to fine-tune, retrieve, or prompt.

A payback volume, plus the retraining cadence that resets it.

A break-even task volume for the narrow model against the general one.

The short version

What are the four findings that change a cost decision?

Open the section

4 source sections, in the order this decision needs them.

The sections below are the report's own material, composed here. Nothing is rewritten for this audience; only the order and the framing are.

Step 01 · orientation The short version What are the four findings that change a cost decision?

The cost frame this decision sits inside.

Open this section on its own

The sticker price of a token is not the cost of an answer. In August 2026, the same standard API workload (100K input + 20K output tokens) costs anywhere from $0.018 on Gemini Flash-Lite to $2.00 on Claude Fable 5 — a 111x price range across comparable models. Two corrections cut the other way: measured token-efficiency variance means the same input produces 2.65x+ more output tokens on some models, and cheap-token models can cost more per successful task via verbosity and retry loops — so the posted range is an upper bound on effective dispersion, and it still understates the real spread: adjusted for quality using the Artificial Analysis Intelligence Index, the cost-per-quality- point spreads by 47x. Read that 47x as order-of-magnitude evidence of non-linear pricing, not a precise ratio: the index is an editorially-weighted nine-benchmark composite (text/English-centric; agentic-weighted since v4.1), and version overhauls move scores by ~23 points — as large as the frontier spread itself. The marginal cost of capability accelerates sharply above Intelligence Index 50.

Running your own GPUs is a crossover problem, not a preference. Specialist GPU clouds (“neoclouds” like RunPod, Lambda, Together AI) charge 50-70% less than hyperscalers for the same H100: a median of $3.99/GPU-hour vs $7.89. The TCO model — now corrected for procurement scope — shows the answer splits by who’s buying: a lean operator adding GPUs to existing infrastructure (~$30K/GPU street price) breaks even against hyperscaler on-demand at ~21% utilization and beats median neocloud pricing above ~37%; an enterprise node-loaded buyer (~$94K/GPU all-in per io.net’s measured 3-year figure) breaks even vs hyperscaler on-demand only at ~57%, essentially never beats median neocloud pricing, and never beats cheap neoclouds or spot. Reserved commitments beat node-loaded on-prem at every utilization. Utilization assumption AND procurement scope together dominate the decision.

The control premium is real, and buyers pay it knowingly. Independent cost itemization puts true self-hosting at 3–5× the pure GPU price — matching the node-loaded math. For a compliance-driven subset, the premium is not optional: sovereign cloud IaaS spending is forecast at $80B in 2026 (+35.6%), Italy’s cloud market grew 20% YoY explicitly on sovereign-data demand, and the EU’s Technological Sovereignty Package (June 2026) extends data-governance obligations to AI providers. The other evidenced motives: guaranteed capacity with no rate limits or peak contention, latency for on-network workloads, vendor independence as a hedge against the list-price and subsidy moves forecast above, and very-high sustained utilization. In practice the control premium runs roughly 2–4× effective cost vs neoclouds — cost-optimal is not decision-optimal when compliance or capacity certainty is a hard constraint.

A significant fraction of current AI spend is subsidized — through provider credits, free tiers, academic programs, and below-cost pricing funded by venture capital — though the exact fraction is opaque. The effective cost of AI is materially higher than what most teams currently pay, and the subsidy distorts the build-vs-buy decision by making API inference appear cheaper than its true cost.

Step 02 · decision If you're evaluating fine-tuning vs retrieval When does the narrow model win, and by how much?

Where the compute cost stops and the real cost starts.

Open this section on its own

Three verified findings from the fine-tuning notebook, all pointing the same direction. PEFT (LoRA/QLoRA) makes the compute almost free — a 70B QLoRA run costs tens of dollars on a single H100 vs 8×H100 hours for full fine-tuning. Compute is nonetheless only 2–3% of a real programme: dataset curation, evaluation, and MLOps dominate, and two conditions widen that gap — fine-tuning causes unpredictable safety drift even on clean data (budget recurring safety evals per refit), and every base-model deprecation forces a refit, making compute a subscription rather than a purchase. On the build decision itself: RAG and fine-tuning are substitutes for knowledge-heavy tasks (RAG is updateable and cheaper to iterate) but complements for behavior-heavy tasks — fine-tune the behavior, retrieve the knowledge.

Step 03 · evidence Inference economics — five verified claims What does a token actually cost, and how far does effective cost diverge from the posted price?

The per-token baseline the narrow model has to beat in production, not in training.

Open this section on its own

The inference-economics notebook has 17 references and 5 verified claims:

  1. C1 (verified, 3 independent grade-A sources): The same API workload costs $0.018 to $2.00 across major LLMs — a 111x price range driven by the 21x blended-price spread in the frontier basket.

  2. C2 (verified, 3 independent grade-A sources): Quality-adjusted cost diverges from token price by 47x across the frontier basket — far more than the 2x the original hypothesis predicted. DeepSeek V4-Flash at $2/index-point vs Claude Fable 5 at $94/index-point.

  3. C3 (verified, 3 independent grade-A sources): The price-per-quality frontier is non-linear: models in the 40-50 Intelligence Index range offer 10-47x better cost-per-index-point than models in the 55-60 range. The marginal cost of capability accelerates above index 50.

  4. C4 (verified, 1 grade-A source): Closed-weight flagships carry a 10% premium over open-weight equivalents in blended price ($4.47 vs $4.08/Mtok), but the premium concentrates at the frontier, not uniformly across tiers.

Step 04 · boundary Evidence boundaries What does this evidence not establish?

What the fine-tuning evidence does not settle.

Open this section on its own
  • All prices are evidence of posted rates on 2026-08-20, not durable claims. Provider pricing changes weekly.
  • The quality-adjusted cost analysis uses the Artificial Analysis Intelligence Index, which is a composite benchmark score — it does not predict performance on a specific workload. The index is also an editorially-weighted nine-benchmark composite (text/English-centric, agentic-shifted since v4.1); version overhauls have moved scores by ~23 points, so cost-per-index-point figures are order-of-magnitude evidence, not precise rankings.
  • The GPU pricing data is compiled from independent sources including two grade-A direct captures (getdeploying.com, spendark.com) and a compiled evidence file; both GPU claims are verified.
  • The subsidy analysis is structural (the subsidy exists, it distorts decisions) and now has a cost-side quantitative anchor: frontier-lab list pricing runs below sustainable infrastructure return (OpenAI gross margin 33% in 2025 vs its own 46% forecast, inference costs ~$8.4B and 4x YoY; Anthropic 40%) — posted API prices remain effectively investor-subsidized. The user-side subsidized fraction is still unmeasured.
  • The forecast bets above are dated, binary, probability-bearing judgments over a short record — not calibrated forecasts. FC4 and FC6 have named missing baselines; their probabilities are structured priors.
  • Recheck by: 2026-11-20, or after a material model release, pricing change, or hardware generation shift.

Three questions that keep this route honest.

  1. Check 01

    Have you costed data preparation and evaluation, or only GPU hours?

  2. Check 02

    Is the task knowledge-heavy or behavior-heavy — substitutes or complements?

  3. Check 03

    How often does the underlying knowledge change, and what does that do to payback?

What is still missing

A payback calculator that takes retraining cadence as an input.

Read the whole argument

Routes are a reading order, not a substitute for the report. The full synthesis carries the update log and the complete evidence boundary.