World Labs releases Atlas for controllable 1440p, one-minute 3D world generation
- Atlas takes one or more images plus direct geometric camera movement inputs and generates new views, with output of up to one minute at 1440p.
- World Labs says its model reconstructs scenes from one to dozens of images without special capture equipment, and can accept more than 100 inputs.
- Atlas is trained from scratch on text, images, video, and 3D data; World Labs says it anchors each input to a position in 3D space through shared "spatial context."
- The model produces native 3D outputs including point clouds and 3D Gaussian splats because it processes depth alongside RGB data.
- World Labs claims Atlas outperforms specialized 3D models with two or three images, including generating aerial views from ground-level photos in a Stanford Main Quad demo.