World Labs launches Atlas, a shared model for controlled video, 3D reconstruction and robot views
- World Labs launched Atlas on September 1 as one model for controlled video generation, 3D scene reconstruction, and simulated camera views for robots.
- Atlas is a multimodal autoregressive diffusion transformer pretrained from scratch that accepts text, images, camera poses, and 3D depth maps.
- The model combines inputs into a shared spatial context that grounds each image at a position in three-dimensional space.
- Atlas takes a camera path as native input, giving users direct control over generated-view position and angle rather than relying only on text instructions such as pan or crane.