Black Forest Labs Launches FLUX 3, a Unified Model for Video, Audio, and Robot Control
AI video model releases
Black Forest Labs Launches FLUX 3, a Unified Model for Video, Audio, and Robot Control
Black Forest Labs launched FLUX 3 on July 23, 2026, training one set of weights jointly on images, video, and audio, then extending the architecture to predict robot actions.
FLUX 3 Video, the only component in broad early access, generates clips up to 20 seconds at 720p with audio produced in the same generation pass, unlike rivals Runway, Luma, and Kling that add audio as a separate pass.
FLUX-mimic, a robotics model built on the FLUX 3 architecture, is already running on production lines at Audi.
In BFL's internal human evaluations on 10-second 720p clips, FLUX 3 was preferred 93% of the time over Luma Ray 3.2, 77% over Runway Gen-4.5, 69% over Grok Imagine Video, 60% over Kling v3 Pro, and 52% over Seedance 2.0 and Gemini Omni Flash.
The model uses a method called Self-Flow, published by BFL in March 2026, and video prediction accounts for over 95% of training compute while audio takes up less than 0.5% of tokens in a 720p generation.