decision-sequence · section 11 of 13

Setting your 2027 AI budget — the six-decision sequence

I'm writing a 2027 AI budget right now — what does this evidence tell me to do, in what order?

Six sequential decisions: assume unit cost falls while total spend rises; treat subsidies as expiring and budget at list price; decide capacity ownership by procurement scope; keep the routing stack portable; budget evaluation explicitly; write reopen conditions into the budget itself.

Every number is evidence-dated 2026-08-20/21 and carries its stated conditions; the forecast bets are probability-bearing judgments, not calibrated forecasts.

Report heading Setting your 2027 AI budget — the six-decision sequence, under What this means for decisions, in The real cost of AI: August 2026.

The section itself

If you are writing a 2027 AI budget between now and January, the evidence in this report maps onto six sequential decisions. Each names the job, what the verified evidence says, and what would change the answer.

1. Set the envelope: assume unit cost falls, total spend rises. Fixed- quality inference is still getting cheaper (1.5–5x/year and decelerating), but the Jevons bet (FC3, p=0.65) says volume grows faster than price falls. Budget the unit line shrinking and the volume line growing; a budget that assumes both flat will be spent differently than written by Q2.

2. Choose your commitment structure: treat subsidies as expiring. The cost-side anchor is verified — frontier list prices run below sustainable infrastructure return (OpenAI’s gross margin fell to 33% against its own 46% forecast). The subsidy-contraction bet (FC4, p=0.50) says effective cost for subsidized teams rises 1.5–3x over the window. Practical rule: build your base budget at list price with no credits, then treat credits as a discount to be won, not a floor to stand on. Audit which of your workloads sit on zero-price channels (:free variants, launch access) — that channel bites first if contraction happens.

3. Decide capacity ownership by procurement scope, not by preference. The two-regime answer: lean buyers adding GPUs to existing infrastructure break even vs hyperscaler on-demand at ~21% utilization and beat median neoclouds above ~37%; enterprise node-loaded buyers (~$94K/GPU all-in) need ~57% and effectively never beat neoclouds or spot. If you’re buying anyway for sovereignty, compliance, or guaranteed capacity, price it as insurance — roughly 2–4x effective cost — because regulatory sovereignty demand is measurably expanding ($80B sovereign IaaS forecast for 2026, EU Sovereignty Package in force).

4. Structure the routing/middleware stack for exit. The layer charges a thin flat fee (~5.5%), delivers savings an order of magnitude larger, and is decommoditizing under your feet — open-source gateways replicate it free, switching costs are near zero, and consolidation has begun (Stripe/OpenRouter). Use routers for transport; keep model choice, evaluation thresholds, and prompt architecture portable. The fit problem — knowing which model clears your bar, and paying when it doesn’t — stays yours no matter what you route through.

5. Budget evaluation explicitly. The routing-savings evidence is quality- conditioned (58% realized at 91% acceptance in the measured pilot), and the quality index behind any comparison is editorially weighted. Your acceptance threshold is a business decision that needs its own eval budget — the line item most teams forget, and the one that determines whether the routing savings are real.

6. Write the reopen conditions into the budget itself. Every number above is evidence-dated 2026-08-20/21. Name what reopens each line: a material repricing or model release, subsidy program changes, :free channel cap moves, or the quarterly capex map. A budget with reopen conditions survives contact with the market; one without them gets defended after it’s wrong.

Worked illustration — a 20-engineer product team

Illustrative arithmetic from the custody rates above (not a claim about any team; substitute your own volumes). Suppose the team spent ~$96K on inference in 2026 at blended list rates, has no compliance constraints, and is writing its 2027 budget:

  • Envelope: unit costs fall 2–4x for fixed quality (FC1/FC2 trajectory); volume grows faster (FC3). Line item: $60–80K at list for more usage than 2026 — not $96K for the same usage.
  • Commitments: no multi-year lock-ins; credits booked as upside, not base. Exposure if :free channels or startup credits vanish: priced at zero.
  • Capacity ownership: at this scale, procurement scope is enterprise-grade if bought new (~$94K/GPU all-in) — ownership loses to neoclouds at every utilization below ~99%. Answer: none. Revisit only if sustained utilization projection exceeds ~70% and hardware is already sunk.
  • Routing stack: one flat-fee aggregator or a self-hosted gateway behind an abstraction boundary; model choice re-rankable monthly. Budget line: ~$0 for transport, ~$6–10K for evaluation (the acceptance-threshold work).
  • Evaluation budget: the acceptance threshold from the measured pilot is the difference between 58% savings and silent quality loss. Fund it as a first-class line, not spare time.
  • Reopen conditions written in: repricing >20% by any incumbent; FC4 evidence (credit-program contraction); Rubin shipment data moving GPU rental prices >25%.

The pattern generalizes: at small-to-mid scale the budget is mostly API line items plus an eval budget; ownership enters only through compliance or very high utilization, and middleware should never be a lock-in.

Worked illustration — the compliance-constrained variant

Same team size, but the workload processes regulated personal data that cannot leave a defined jurisdiction, and projected utilization is high (~70% sustained) because the inference serves a core product loop. Now the decision tree changes shape:

  • The API-first answer is not disqualified by price — it’s disqualified by constraint. If no provider region satisfies the residency requirement, the comparison starts between sovereign-hosted and owned.
  • Sovereign/regional hosting: expect regional-provider rates at or above median neocloud pricing ($4–6/GPU-hour effective in custody examples) plus residency premiums. At 70% utilization that’s roughly $1,900–2,900/GPU-month.
  • Owned, node-loaded: ~$2,900/GPU-month fully loaded at 100% utilization (io.net basis), scaling down modestly with idle power. At ~70% utilization the two lanes reach parity — which is exactly why sovereignty-era buyers are choosing ownership at utilizations the pure-cost model calls marginal.
  • What tips it: guaranteed capacity (no contention at peak), auditability of every token path, and freedom from list-price moves (FC4’s upside case makes this worth more, not less).
  • What to budget anyway: the safety-drift and refit conditions apply twice over — owned models still deprecate, and compliance regimes require documented evaluation per refit.

The general rule for constrained buyers: the control premium is not a cost overrun, it’s the price of the constraint — and at high sustained utilization it approaches zero. The teams overpaying are the ones who buy the premium without having the constraint.

Who uses this section, and for what

Developers and ML engineers

The six-decision sequence for writing a 2027 AI budget — envelope, commitments, capacity ownership, routing portability, eval budget, reopen conditions.

Founders and product teams

The six-decision sequence for writing a 2027 AI budget — envelope, commitments, capacity ownership, routing portability, eval budget, reopen conditions.

Reader questions that route through here

In sequence