World Labs says Atlas generates camera-controlled views and 3D scenes from sparse images
- World Labs released Atlas, a multimodal world model that takes text, images, video, and camera poses and generates camera-conditioned views, reconstructs scenes, and supports simulation.
- Atlas uses novel view prediction: given a few views of a scene and their spatial relationships, it predicts what the scene looks like from a requested new camera position.
- The team says Atlas can reconstruct a scene from as few as three camera inputs, compared with the hundreds or thousands of images often used for conventional capture, and demonstrated bullet-time-style shots from three iPhones.
- World Labs plans to add dynamics, editing, and interaction so Atlas can model how generated environments change after actions, with robotics simulation named as a target use.
- Fei-Fei Li compares novel view prediction with next-token prediction in LLMs and describes it as a possible foundational primitive for spatial intelligence and AGI.