Black Forest Labs launches Flux 3, a video-and-audio model it also aims at robot control via Flux-mimic
AI video model releases
Black Forest Labs launches Flux 3, a video-and-audio model it also aims at robot control via Flux-mimic
Black Forest Labs released FLUX 3, its first multimodal model generating up to 20 seconds of video with synced dialogue, sound effects, and background music from a single training system covering images, video, and audio.
In human preference tests, FLUX 3 beat Runway Gen-4.5 77% of the time and Luma Ray 3.2 93% of the time, and it won 52% of comparisons against Google Gemini Omni and Seedance (company-reported, not a standardized benchmark).
The company paired FLUX 3 with Zurich startup Mimic Robotics to build FLUX-mimic, which uses a lightweight decoder on top of FLUX 3's video-prediction engine to translate video motion into robot actions, with system response time around 101 milliseconds.
Audi is testing FLUX-mimic on a production line for door seal installation, a soft-material assembly task that existing automated equipment has struggled with.
Access is limited for now: video and robot-action features are API-only for select partners, image generation is coming in following weeks, and a locally usable public-weights FLUX Dev version won't ship until the second half of 2026.