decision · section 08 of 13

The routing layer itself — thin fees, thin moats

Is a routing layer worth its take rate, and is it a business or a feature?

Routing savings are real but quality-conditioned, and the take rate charged for them is thin enough that the layer is a feature more than a moat.

Reported savings ranges come from vendor-adjacent pilots; the acceptance threshold that produced them travels with the number.

Report heading The routing layer itself — thin fees, thin moats, under What this means for decisions, in The real cost of AI: August 2026.

The section itself

The middleware that performs the routing is not where the money is, and may not be where it stays. OpenRouter — the category leader, processing over $100M in annualized inference spend and more than a quadrillion tokens/year by mid-2026 — charges a flat 5.5% on credit purchases with token prices passed through at provider list rates. Its fee is roughly one-tenth the size of the savings its own category claims to deliver. The layer is also structurally commoditizing: open-source gateways (LiteLLM) replicate the unified-API function free forever, 10+ commercial competitors ship the same feature set, and the standardized OpenAI-compatible interface makes switching cost nearly zero. Stripe’s $7.5B acquisition of OpenRouter (August 2026) reads as the category’s sustainability answer: routing survives as payments infrastructure, not as a standalone margin business.

Who absorbs the cost of model fit? Not the router. The 5.5% fee is the smallest line in the fit-cost stack; the real costs of matching models to workloads — evaluation effort to pick the threshold model, retry and quality-mismatch costs when the cheap model fails, re-integration each time a current model is deprecated — are absorbed by the developer. Routing-as-a- service removes the transport problem, not the fit problem.

Exit checklist — how to keep switching cost near zero. Since the layer’s standardization is what makes switching cheap, portability is a property you keep or lose by construction. Six checks, all cheap to maintain from day one:

  1. Speak only the OpenAI-compatible schema at your client boundary — no provider-native SDK types leaking past it.
  2. Keep the eval harness router-independent: acceptance thresholds measured against model outputs directly, so re-routing never invalidates your bar.
  3. Export usage/cost logs to storage you own — billing data is the lever in any renegotiation.
  4. Avoid provider-exclusive parameters (fine-grained logit control, native tool dialects) unless the dependency is worth a migration.
  5. Re-run the quality threshold quarterly against the current frontier basket — the non-linear frontier means last quarter’s “cheap model that clears” may no longer be on it.
  6. Price the exit annually: hours to swap gateways × loaded rate. If that number grows, the middleware has become infrastructure — renegotiate or migrate before it becomes both.

Who uses this section, and for what

Reader questions that route through here

In sequence