Black Forest Labs Launches FLUX 3, a Multimodal Model Spanning Image, Video, Audio and Robot Action
AI video model releases
Black Forest Labs Launches FLUX 3, a Multimodal Model Spanning Image, Video, Audio and Robot Action
Black Forest Labs, based in Freiburg, Germany, announced FLUX 3, a multimodal frontier model jointly trained on image, video, and audio within a single unified architecture, with an extension for action prediction.
FLUX 3 builds on Self-Flow, the company's method for aligning multimodal generation and understanding in one architecture, and testing showed the same architecture can extend to action prediction for robotics without losing video generation quality.
The model ships in four variants: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and FLUX 3 Dev, with FLUX 3 Video already leading in early evaluations against frontier video models on facial expressions, sound-to-event association, and multilingual handling.
Canva, Burda, Magnific (formerly Freepik), Krea, and Picsart are already testing FLUX 3, and the FLUX family already powers features in Adobe Photoshop and Nous Research's Hermes Agent.
CEO Robin Rombach said joint training across modalities lets each one reinforce the others, since audio conveys timing and physical events while video teaches dynamics that still images cannot capture.