decision · section 05 of 10

If you run ML infrastructure

Rent, reserve, or own — at what utilization and under whose procurement scope?

The answer is a utilization threshold, not a preference, and the threshold moves with whether you are adding GPUs to existing infrastructure or buying loaded nodes.

Compliance, capacity certainty, and latency constraints can make the cost-optimal answer the wrong one; the control premium runs roughly 2–4x.

The section itself

The GPU market is split: neoclouds at $3-4/GPU-hour, hyperscalers at $7-12. If you’re paying hyperscaler on-demand for sustained inference, you are likely overpaying by 2x. Whether owning beats renting now has a two-part answer. Procurement scope first: a lean team adding GPUs to existing infrastructure (~$30K/GPU) breaks even vs hyperscaler on-demand at ~21% utilization and beats median neocloud pricing above ~37% — but an enterprise node-loaded buyer (~$94K/GPU all-in) needs ~57% just against hyperscaler on-demand, essentially never beats median neoclouds, and never beats cheap neoclouds or spot. Utilization second: above ~70% sustained, on-prem wins on raw cost in the lean regime. And if you’re buying for sovereignty, compliance, or guaranteed capacity, you’re paying a 2–4x control premium on purpose — price it as insurance, not as infrastructure.

Who uses this section, and for what

Reader questions that route through here

In sequence