Summary
On July 30, 2026, Vercel made Inkling Small from Thinking Machines available on its AI Gateway. Inkling Small is an open-weights multimodal Mixture-of-Experts model with a 1M-token context window and adjustable thinking effort, accessible through the gateway's standard interface.
What changed
Inkling Small, an open-weights multimodal MoE model from Thinking Machines, launched on Vercel AI Gateway. It offers a 1M-token context window and a configurable thinking-effort control, letting developers trade latency against reasoning depth per request.
Why it matters
An open-weights MoE with a 1M-token context gives teams a long-context, tunable-reasoning option they can also self-host, hedging against closed-model lock-in while still using Vercel's gateway for convenience. Its arrival marks Thinking Machines showing up in mainstream developer distribution, not just research channels.
Evidence excerpt
Inkling Small from Thinking Machines is now available on AI Gateway - an open-weights multimodal Mixture-of-Experts model featuring a 1M token context window and adjustable thinking effort.