World Labs unveils Atlas, a multimodal world model that generates controllable 1440p videos up to one minute
- World Labs introduced Atlas, an omni world model pretrained from scratch to work natively with text, images, video, and 3D inputs in a shared spatial context.
- Atlas generates camera-controlled video from one to six reference images, with manually designed camera paths, at up to 1440p for one minute.
- The model accepts precise camera geometry as a native input, generating new views intended to preserve the content and 3D geometry of reference images while filling in unseen areas.
- Atlas reconstructs scenes from one to dozens of images, producing both novel-view image frames and explicit 3D outputs; World Labs says it outperforms specialized 3D reconstruction models.
- Atlas also models space and time from video for reframing effects and Real-to-Sim robotics workflows, and will power future versions of Marble.