NVIDIA releases Cosmos 3 open multimodal world model for physical AI
- Cosmos 3 is NVIDIA's fully open omnimodel for physical AI, with native understanding and generation of text, images, video, ambient sound and actions.
- Its mixture-of-transformers architecture combines a reasoning transformer with an expert generation transformer to model object interactions, motion and spatiotemporal relationships before producing video and action trajectories.
- NVIDIA says Cosmos 3 was trained on billions of multimodal samples spanning text, images, video, sound and action trajectories, targeting synthetic-data generation and physical AI policy models.
- The lineup includes Cosmos 3 Super for post-training robotics and autonomous-vehicle models, Cosmos 3 Nano for sub-second video and action reasoning, and Cosmos 3 Edge for real-time edge inference, which is coming soon.
- NVIDIA formed the Cosmos Coalition with Agile Robots, Black Forest Labs, Generalist, LTX, Runway and Skild AI to contribute models, research and evaluation techniques for open world models.