Black Forest Labs turns FLUX 3 into a video-action model for robots with FLUX-mimic
AI video model releases
Black Forest Labs turns FLUX 3 into a video-action model for robots with FLUX-mimic
Black Forest Labs built FLUX-mimic, a video-action model on top of the FLUX 3 backbone in partnership with robotics startup mimic robotics, and deployed it on robots tested at Audi.
FLUX 3 is trained jointly on images, video and audio from the start, with video prediction accounting for over 95% of total training compute cost.
In a large training run, adding action prediction to the model caused human quality ratings on text-to-video and image-to-video to drop up to 10%, but the model recovered full video quality after 3500 steps while also predicting actions.
FLUX-mimic decodes robot actions from intermediate features of FLUX 3's video prediction path using a lightweight action decoder, a method the companies say was pioneered in mimic-video.
Black Forest Labs frames this as proof that one foundation model can serve both content generation and Physical AI (robot control) without needing separate models for each.