Summary
On September 1, 2026, Fei-Fei Li's World Labs unveiled Atlas, an omni world model pretrained from scratch to natively operate on text, images, video, camera poses, and 3D depth maps in a shared spatial context. Atlas can reconstruct, generate, and simulate 3D environments at 1440p and up to a minute long, including expansive scenes with precise camera control built from a single image, launching in early access.
What changed
World Labs launched Atlas in early access, a natively multimodal spatial model that grounds text, images, video, camera poses, and depth maps in a shared 3D context to generate and simulate camera-controlled 1440p environments up to one minute, building on its earlier Marble model and World API.
Why it matters
Spatial, camera-controllable world models push generative AI beyond flat text and 2D media toward simulated 3D environments useful for robotics, gaming, design, and embodied agents. A single-image-to-navigable-3D capability lowers the cost of building spatial training and simulation data, an emerging layer of AI infrastructure distinct from language models.
Evidence excerpt
Atlas is an omni model pretrained from scratch to natively operate on text, images, video, and 3D... it can generate 3D images and videos at 1440p resolution and up to one minute in length, with precise camera control, enabling them to be viewed from any angle.