Black Forest Labs launches FLUX 3, a multimodal flow model generating video, audio, and images together
AI video model releases
Black Forest Labs launches FLUX 3, a multimodal flow model generating video, audio, and images together
Black Forest Labs released FLUX 3, a multimodal foundation model trained jointly on images, video, and audio using a single unified architecture built on their Self-Flow method.
FLUX 3 Video generates videos up to 20 seconds long with native synchronized audio in a single pass, supporting text-to-video, image-to-video, video-to-video, keyframe-to-video, and multilingual dialogue.
In preliminary preference evaluations, FLUX 3 beat Runway Gen-4.5 in 77% of comparisons, Luma Ray 3.2 in 93%, Grok Imagine Video in up to 69%, and Kling v3 Pro in 60%.
The model can chain individual clips agentically into multi-shot sequences lasting several minutes while keeping character appearance consistent across scenes using visual references.
FLUX 3 Video is now available in Early Access, alongside improved image synthesis and editing capabilities inherited from the same unified model.