Black Forest Labs Launches FLUX 3, a Joint Image-Video-Audio-Action Model, With FLUX-mimic for Robotics
AI video model releases
Black Forest Labs Launches FLUX 3, a Joint Image-Video-Audio-Action Model, With FLUX-mimic for Robotics
Black Forest Labs released FLUX 3, a multimodal model jointly trained on images, video, and audio in one architecture, with an extension for action prediction.
FLUX 3 builds on the company's Self-Flow method and ships in four variants: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and FLUX 3 Dev.
Early evaluations claim FLUX 3 Video leads frontier video models on facial expression capture, sound-to-event association, and multilingual generation, though it remains in development.
Canva, Burda, Magnific, Krea, and Picsart are already testing FLUX 3 ahead of wider release.
Black Forest Labs and mimic robotics unveiled FLUX-mimic, a video-action model built on FLUX 3 that Audi is testing, which can fine-tune for a manipulation task with as little as 30 minutes of robot data versus 30-plus hours previously needed.