Black Forest Labs launches FLUX 3, a multimodal video, image and audio model with 20-second native-audio video generation
AI video model releases
Black Forest Labs launches FLUX 3, a multimodal video, image and audio model with 20-second native-audio video generation
Black Forest Labs released FLUX 3 in Early Access as a multimodal foundation model trained jointly on images, video, and audio using an architecture built on their Self-Flow approach.
FLUX 3 Video generates up to 20 seconds of video with native audio in a single generation, and supports text-to-video, image-to-video, video-to-video, keyframe-to-video, and multilingual dialogue.
In early preference evaluations on 10-second 720p text-to-video clips, FLUX 3 beat Runway Gen-4.5 in 77% of comparisons and Luma Ray 3.2 in 93%, but only 52% against Seedance 2.0 and Gemini Omni Flash.
Black Forest Labs plans to release an open-weight multimodal backbone called FLUX 3 Dev for content creation and action prediction, alongside a research partnership with mimic robotics on action prediction (FLUX-mimic and FLUX 3 Action).
Hacker News commenters flagged the blog post's writing as "LLM slop" and criticized the launch for showing almost no video examples and few human subjects despite the 20-second claim.