Summary
On July 21, 2026 Vercel added service tiers to its AI Gateway, letting developers trade latency and cost on a per-request basis with automatic billing adjustments so the same route can be tuned for speed or savings.
What changed
Vercel AI Gateway introduced service tiers that let each request opt into different latency/cost tradeoffs, with billing adjusting automatically.
Why it matters
Per-request service tiers give teams a knob to control inference spend without changing code paths, reflecting how model routing is maturing into a cost-management layer. It strengthens the AI Gateway as a place to govern spend across many models, not just proxy calls.
Evidence excerpt
Service tiers now available on AI Gateway allow trading latency and cost per-request with automatic billing adjustments.