Black Forest Labs launches FLUX 3, one model that generates images, video, audio, and robot actions · cho.sh
Black Forest Labs launches FLUX 3, one model that generates images, video, audio, and robot actions
AI video model releases
Black Forest Labs launches FLUX 3, one model that generates images, video, audio, and robot actions
Black Forest Labs released FLUX 3 on July 23, training one backbone on images, video, and audio together, then extending it to predict robot actions instead of building separate models for each output type
FLUX 3 Video generates clips up to 20 seconds long in a single pass with native audio (dialogue, sound effects, ambient noise), matching the length of the shut-down Sora and beating most current rivals on duration
In BFL's own preliminary tests on 10-second 720p clips, reviewers preferred FLUX 3 over Luma Ray 3.2 in 93% of comparisons and Runway Gen-4.5 in 77%, but only 52% over ByteDance Seedance 2.0 and Google Gemini Omni Flash
FLUX 3 Action is already being tested on a real Audi production line through robotics startup mimic, while FLUX 3 Video remains gated early access and the open-weight FLUX 3 Dev backbone is not due until later in 2026
The research idea behind the model, called Self-Flow, trains one shared representation across image, video, and audio so that generated sound must match motion and motion must obey physical constraints