If you are writing a 2027 AI budget between now and January, the evidence in
this report maps onto six sequential decisions. Each names the job, what the
verified evidence says, and what would change the answer.
1. Set the envelope: assume unit cost falls, total spend rises. Fixed-
quality inference is still getting cheaper (1.5–5x/year and decelerating), but
the Jevons bet (FC3, p=0.65) says volume grows faster than price falls. Budget
the unit line shrinking and the volume line growing; a budget that assumes
both flat will be spent differently than written by Q2.
2. Choose your commitment structure: treat subsidies as expiring. The
cost-side anchor is verified — frontier list prices run below sustainable
infrastructure return (OpenAI’s gross margin fell to 33% against its own 46%
forecast). The subsidy-contraction bet (FC4, p=0.50) says effective cost for
subsidized teams rises 1.5–3x over the window. Practical rule: build your
base budget at list price with no credits, then treat credits as a discount
to be won, not a floor to stand on. Audit which of your workloads sit on
zero-price channels (:free variants, launch access) — that channel bites
first if contraction happens.
3. Decide capacity ownership by procurement scope, not by preference. The
two-regime answer: lean buyers adding GPUs to existing infrastructure break
even vs hyperscaler on-demand at ~21% utilization and beat median neoclouds
above ~37%; enterprise node-loaded buyers (~$94K/GPU all-in) need ~57% and
effectively never beat neoclouds or spot. If you’re buying anyway for
sovereignty, compliance, or guaranteed capacity, price it as insurance —
roughly 2–4x effective cost — because regulatory sovereignty demand is
measurably expanding ($80B sovereign IaaS forecast for 2026, EU Sovereignty
Package in force).
4. Structure the routing/middleware stack for exit. The layer charges a
thin flat fee (~5.5%), delivers savings an order of magnitude larger, and is
decommoditizing under your feet — open-source gateways replicate it free,
switching costs are near zero, and consolidation has begun (Stripe/OpenRouter).
Use routers for transport; keep model choice, evaluation thresholds, and prompt
architecture portable. The fit problem — knowing which model clears your bar,
and paying when it doesn’t — stays yours no matter what you route through.
5. Budget evaluation explicitly. The routing-savings evidence is quality-
conditioned (58% realized at 91% acceptance in the measured pilot), and the
quality index behind any comparison is editorially weighted. Your acceptance
threshold is a business decision that needs its own eval budget — the line item
most teams forget, and the one that determines whether the routing savings are
real.
6. Write the reopen conditions into the budget itself. Every number above
is evidence-dated 2026-08-20/21. Name what reopens each line: a material repricing
or model release, subsidy program changes, :free channel cap moves, or the
quarterly capex map. A budget with reopen conditions survives contact with the
market; one without them gets defended after it’s wrong.
Worked illustration — a 20-engineer product team
Illustrative arithmetic from the custody rates above (not a claim about any
team; substitute your own volumes). Suppose the team spent ~$96K on inference
in 2026 at blended list rates, has no compliance constraints, and is writing
its 2027 budget:
- Envelope: unit costs fall 2–4x for fixed quality (FC1/FC2 trajectory);
volume grows faster (FC3). Line item: $60–80K at list for more usage than
2026 — not $96K for the same usage.
- Commitments: no multi-year lock-ins; credits booked as upside, not base.
Exposure if :free channels or startup credits vanish: priced at zero.
- Capacity ownership: at this scale, procurement scope is enterprise-grade
if bought new (~$94K/GPU all-in) — ownership loses to neoclouds at every
utilization below ~99%. Answer: none. Revisit only if sustained utilization
projection exceeds ~70% and hardware is already sunk.
- Routing stack: one flat-fee aggregator or a self-hosted gateway behind
an abstraction boundary; model choice re-rankable monthly. Budget line: ~$0
for transport, ~$6–10K for evaluation (the acceptance-threshold work).
- Evaluation budget: the acceptance threshold from the measured pilot is
the difference between 58% savings and silent quality loss. Fund it as a
first-class line, not spare time.
- Reopen conditions written in: repricing >20% by any incumbent; FC4
evidence (credit-program contraction); Rubin shipment data moving GPU rental
prices >25%.
The pattern generalizes: at small-to-mid scale the budget is mostly API line
items plus an eval budget; ownership enters only through compliance or very
high utilization, and middleware should never be a lock-in.
Worked illustration — the compliance-constrained variant
Same team size, but the workload processes regulated personal data that cannot
leave a defined jurisdiction, and projected utilization is high (~70%
sustained) because the inference serves a core product loop. Now the decision
tree changes shape:
- The API-first answer is not disqualified by price — it’s disqualified by
constraint. If no provider region satisfies the residency requirement, the
comparison starts between sovereign-hosted and owned.
- Sovereign/regional hosting: expect regional-provider rates at or above
median neocloud pricing ($4–6/GPU-hour effective in custody examples) plus
residency premiums. At 70% utilization that’s roughly $1,900–2,900/GPU-month.
- Owned, node-loaded: ~$2,900/GPU-month fully loaded at 100% utilization
(io.net basis), scaling down modestly with idle power. At ~70% utilization
the two lanes reach parity — which is exactly why sovereignty-era buyers are
choosing ownership at utilizations the pure-cost model calls marginal.
- What tips it: guaranteed capacity (no contention at peak), auditability
of every token path, and freedom from list-price moves (FC4’s upside case
makes this worth more, not less).
- What to budget anyway: the safety-drift and refit conditions apply twice
over — owned models still deprecate, and compliance regimes require
documented evaluation per refit.
The general rule for constrained buyers: the control premium is not a cost
overrun, it’s the price of the constraint — and at high sustained utilization
it approaches zero. The teams overpaying are the ones who buy the premium
without having the constraint.