Black Forest Labs debuts FLUX 3, a unified image, video and audio model with early robotics use
AI video model releases
Black Forest Labs debuts FLUX 3, a unified image, video and audio model with early robotics use
Black Forest Labs launched FLUX 3, a multimodal foundation model trained jointly on images, video, and audio, also used for robotic action prediction
FLUX 3 Video generates clips with native audio up to 20 seconds long from text, images, or existing video, and supports continuation, keyframe transitions, multilingual dialogue, typography, and clip chaining
In the company's own preliminary tests, FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen 4.5 in 77%, Grok Imagine Video in 69%, and Kling v3 Pro in 60%, though full methodology is not yet published
The video backbone is adapted for robotics through FLUX mimic, built with mimic robotics, and Audi is testing it for production manipulation tasks that can be fine tuned with as little as 30 minutes of robot data
FLUX 3 Video and FLUX 3 Action are in early access now, with API access, private weights, and an open weight FLUX 3 Dev version planned later this year