Black Forest Labs Turns FLUX 3 Video Model Into Robot Action Model With mimic robotics
- Black Forest Labs and mimic robotics built FLUX-mimic, a video-action model that runs robots by decoding actions from the internal world representation learned by the FLUX 3 multimodal foundation model.
- FLUX 3 is trained jointly on images, video, and audio, with video prediction accounting for over 95% of total training compute, while audio makes up less than 0.5% of tokens in a 720p video with audio.
- When action prediction was added to FLUX 3's training curriculum, human ratings on text-to-video and image-to-video quality dropped by up to 10% before recovering to full prior performance after 3500 steps, showing actions can be learned on the same backbone without permanent capability loss.
- mimic robotics builds the physical robots and deployment stack (already tested at Audi), while BFL supplies the multimodal foundation model, and a lightweight action decoder is trained on intermediate features from FLUX's video prediction path, an approach pioneered in mimic-video.
- Black Forest Labs frames this as extending its roadmap from content generation (images, video, audio) to Physical AI (robot actions) using one shared foundation model rather than building a separate model for robotics.
Hacker News opinions
Nice to see a European startup partnership. Wait, wasn't Flux bought by Meta though?
If Amazon doesn't buy them the CEO should get promoted, not fired lol. Hope it stays European, maybe Mistral buys them or they team up.
BFL is in the same region as BMW and Audi. If Audi sees BFL as the future of factory automation, it'll be hard for anyone to buy them out.
Really interesting: a well trained multimodal video model has a world model trained inside it, and they lifted it out to drive robots. I haven't seen a video lab pivot into a robot lab before, this might be train video model, sell video gen, scale, use scale to train robots, profit. Their hands look like gloves with sensors packed in, unlike Xiaomi's Robotics-1 which needs people wearing gripper gloves to generate training video, BFL's approach looks like it skips that.
This tactic isn't new, Nvidia has demoed robots trained with video generation models and Waymo's been doing it too.
FYI other video labs are also getting into robotics, Luma Labs and Runway both announced physical AI efforts.
The clip around 3:30 where the robot arm took three attempts to reseat the window trim was unnerving, I haven't seen resolving like that before, is this new?
You're out of the loop, Google did something arguably more impressive over a year ago with a VLA bot replacing a tensioned timing belt.
If you want state of the art dexterous manipulation, check out Generalist's Gen-1.
That phrasing about 'less disentangled representations' reads like something only an LLM would write when explaining why less disentangled reps are less useful.
Humans write awkward stuff like that all the time, honestly LLMs are probably less likely to produce clunky phrasing since it's rare in training data.
We're just guessing 'LLM wrote this' based on vibes now, might as well assume everything is AI assisted and judge the writing on its own merit.
We're heading into a crisis where humans have less value unless they own the means of production, dark days ahead.
This tech is amazing but movies are worse than ever, I watch old films with goofy puppets because the storytelling was so much better back then.
That's survivorship bias, you're just not watching all the forgotten bad movies from back then too.