Summary
On July 8, 2026, Mistral released Robostral Navigate, an 8-billion-parameter vision-language model that lets robots follow plain-language navigation instructions using only a single RGB camera—no LiDAR or depth sensors. It reports a 76.6% success rate on the R2R-CE benchmark for unseen environments, beating the best single-camera approach by 9.7 points.
What changed
Mistral launched Robostral Navigate, an 8B vision-language navigation model that outputs pointing coordinates or displacement commands from a single RGB image and a natural-language instruction, generalizing across wheeled, legged, and flying robots.
Why it matters
Camera-only navigation lowers the hardware cost and complexity of deploying autonomous robots, widening who can build embodied agents. It also marks Mistral's first move into physical AI, extending the frontier-model race into robotics where NVIDIA and others are investing.
Evidence excerpt
Robostral Navigate is an 8B vision-language model that lets robots follow natural language instructions through complex spaces with nothing more than a single ordinary RGB camera ... 76.6% success rate on the R2R-CE benchmark.