Summary
On July 23, 2026, Vercel added Ant Group's Ling 3.0 Flash to AI Gateway, a Mixture-of-Experts model with 124B total parameters (about 5.1B active per token), a 256K context window, and thinking and non-thinking modes aimed at token-efficient agentic inference. It is free through August 3, 2026, then billed at provider pricing with no Gateway markup, including for BYOK.
What changed
Ling 3.0 Flash (inclusionai/ling-3.0-flash-free) became available on Vercel AI Gateway. It is a 124B-total-parameter MoE with about 5.1B active per token, a 256K context window, and both thinking and non-thinking modes, positioned for high-frequency agentic workflows, coding agents, document work, and long-context multi-turn interactions. Free through August 3, 2026; afterward provider pricing with no markup and no platform fee, including BYOK.
Why it matters
Open-weight, MoE-based models tuned for cheap agentic inference keep eroding the cost floor for running agents at scale. Availability on AI Gateway with no markup and BYOK gives builders a low-cost, swappable option for high-volume agent loops, reinforcing Vercel's neutral-router positioning as open models take a growing share of gateway traffic.
Evidence excerpt
A Mixture-of-Experts model with 124B total parameters and about 5.1B active per token... built for token-efficient agentic inference at production scale.