Black Forest Labs Launches Flux 3, a Multimodal Model Generating 20-Second Video with Native Audio · cho.sh
Sign in to create alerts.
Black Forest Labs Launches Flux 3, a Multimodal Model Generating 20-Second Video with Native Audio
AI video model releases
Black Forest Labs Launches Flux 3, a Multimodal Model Generating 20-Second Video with Native Audio
Black Forest Labs released Flux 3, a foundation model trained jointly on images, video, and audio, built on the company's Self-Flow architecture that unifies generation and understanding in one model.
Flux 3 generates video clips up to 20 seconds long with native audio that matches sounds to physical events, and it supports text-to-video, image-to-video, video-to-video, and keyframe transitions.
In 10-second, 720p preference tests, Flux 3 beat Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, and edged out Seedance 2.0 and Gemini Omni Flash at 52% each, though BFL says the results are preliminary with no independent verification yet.
The model includes an action component tested as Flux-mimic, a video-action system run on production tasks at Audi with Mimic Robotics, aimed at robotics applications.
BFL plans to release Flux 3 Image in early access within weeks and open-weight access to the multimodal backbone as Flux 3 Dev, while action prediction stays limited to select partners for now.