Summary
On July 15, 2026, Thinking Machines Lab released Inkling, its first open-weights model, under Apache 2.0. Inkling is a multimodal mixture-of-experts model (975B total / ~41B active parameters, 1M-token context) tuned for coding and agentic reasoning, published with full weights on Hugging Face and offered day one as hosted inference on Vercel AI Gateway, Together AI, Fireworks, Modal, Baseten, and Databricks.
What changed
Thinking Machines Lab published Inkling as an Apache 2.0 open-weights multimodal MoE model with controllable thinking effort (an adjustable 0.2-0.99 reasoning-effort parameter) and a lighter Inkling-Small in preview, shipping full weights (including an NVFP4 Blackwell checkpoint) to Hugging Face alongside day-one API availability on Vercel AI Gateway, Together AI, Fireworks, Modal, Baseten, and Databricks.
Why it matters
A lab founded by former OpenAI CTO Mira Murati releasing a frontier-scale open-weights model that is simultaneously downloadable and hosted across most major inference marketplaces collapses the usual lag between a model launch and its distribution, giving builders immediate provider choice, price competition, and a self-host option -- a direct contrast to the closed flagship models from OpenAI, Anthropic, and Google.
Evidence excerpt
Inkling is a mixture-of-experts transformer with 975B total parameters and 41B active ... available via APIs on TogetherAI, Fireworks, Modal, Databricks, and Baseten, with full weights on Hugging Face.