Cost-per-user is three multiplications, and the subsidy question dominates all
of them. First, tokens per active user: your workload’s input/output profile
times usage frequency — measure it, don’t estimate from demos, because token-
efficiency varies 2.65x+ between models on the same input. Second, the rate:
the frontier’s 111x posted range means model choice sets the rate, and caching
(50–90% off repeated context) plus batch (flat 50%) can cut it again — with
the cache-write caveat that one-shot workloads may not benefit. Third, the
subsidy exposure: if any of your cost base rides zero-price channels (:free
variants) or startup credits, your unit economics are temporary by construction
— price the product at list, and treat today’s subsidy as margin, not as your
cost basis.
The failure mode this session’s evidence most warns about: a product whose
gross margin works only while someone else funds the inference. When the
subsidy contracts (FC4), the teams harmed are precisely those who priced
against subsidized rates without a list-price floor in their model.