orientation · section 13 of 13

Context — the vocabulary this report assumes

What do the technical terms in this report actually mean for my decision?

Plain-language definitions of the fourteen terms doing the heaviest work — from tokens and PEFT to the control premium and procurement scope — each tied to where it matters in a budget or procurement decision.

Definitions are decision-oriented simplifications, not exhaustive technical references.

The section itself

Decision-makers shouldn’t need an ML background to use this evidence. Plain-language definitions of the terms doing the heaviest work:

  • Token — the unit of text models read and write (~¾ of a word). All pricing starts here, which is why token counts, not just rates, matter.
  • PEFT / LoRA / QLoRA — techniques that adapt a large model by training a tiny fraction of it. Makes fine-tuning compute cheap enough to run on one GPU — but doesn’t shrink the surrounding programme costs (data, evaluation, refits).
  • Neocloud — a specialist GPU cloud (RunPod, Lambda, Together AI) renting raw GPU-hours at roughly half of hyperscaler prices, typically without the enterprise contracting layer.
  • On-demand / reserved / spot — three rental tiers: pay-per-hour at list; 30–62% off for committing 1–3 years; and 60–91% off with interruption risk (the machine can be taken back mid-job).
  • Prompt caching — paying once to process a repeated context (system prompts, documents), then 50–90% less to re-read it. Cache writes cost 1.25–2× base, so it only pays when context repeats.
  • Batch API — half price on all tokens in exchange for results within 24 hours instead of seconds.
  • Blended price — input and output rates combined into one number using a fixed ratio, so different models can be compared on one axis.
  • Cost per successful task — what this report argues you should actually measure: total tokens (including retries and verbosity) divided by tasks completed acceptably. The only unit that reflects quality differences.
  • Artificial Analysis Intelligence Index — an independent composite score (nine benchmarks, editorially weighted, text/English-centric) used here as the quality axis. Useful for locating the non-linear region; not precise enough to rank models to a decimal.
  • Jevons paradox — the pattern where cheaper units increase total consumption enough that total spend rises. Bet FC3 asserts it applies here.
  • Control premium — the 2–4× effective cost sovereign/compliance buyers knowingly pay for owning infrastructure rather than renting it.
  • Procurement scope — whether “owning” means buying GPUs into existing infrastructure (~$30K/GPU) or procuring complete nodes with networking and financing (~$94K/GPU). The single assumption that swings every ownership crossover band.
  • Utilization — the share of an owned GPU’s time doing paid work. The load-bearing variable in every own-vs-rent comparison.
  • :free variants — aggregator-hosted model endpoints priced at literally $0/token with request caps. Real usage flows through them; they are also the first thing subsidy contraction would remove.

Who uses this section, and for what

Reader questions that route through here

No reader question routes through this section yet.

In sequence