Black Forest Labs launches FLUX 3, a single model generating video, audio and robot actions
AI video model releases
Sign in to create alerts.
Black Forest Labs launches FLUX 3, a single model generating video, audio and robot actions
FLUX 3, released July 23 2026, is Black Forest Labs' first multimodal model generating video, audio, image edits and robot actions from one set of weights, built on the Self-Flow method
FLUX 3 Video produces clips up to 20 seconds long with native audio in a single generation, supporting text-to-video, image-to-video, video-to-video, keyframe-to-video and video-audio continuation
In human-preference tests on 10-second 720p clips with audio, FLUX 3 beat Luma Ray 3.2 93% of the time and Runway Gen-4.5 77%, but only tied Seedance 2.0 and Gemini Omni Flash at 52%
Training compute split shows video prediction used over 95% of compute while audio made up under 0.5% of tokens, meaning sound was cheap to add once motion understanding was established
The same backbone powers FLUX-mimic, a robot control policy running under 80ms on a single RTX 5090, but access is staged: video and action first, image generation weeks later, and open weights (FLUX 3 Dev) last