Flux 3 X Mimic: The Next Generation of Video-Action Models

kensai1 pts0 comments

FLUX 3 x mimic: The Next Generation of Video-Action Models | Black Forest Labs

Contact Sales

Back to blogResearch<br>Models

FLUX 3 x mimic: The Next Generation of Video-Action Models

July 23, 20269 min read

b]:font-medium [&>strong]:font-medium">An early version of FLUX 3, our new multimodal foundation model, is now running on robots. We gave mimic robotics early access to FLUX.3. Their strength in robot learning and deployment, combined with the model's world knowledge and BFL's foundation model expertise, produced FLUX-mimic: the next generation of video-action models.

FLUX<br>b]:font-medium [&>strong]:font-medium">FLUX 1 and FLUX 2 generate images. FLUX 3 expands into multimodality and generates audio-visual content jointly - and, at the same time, provides the foundation of FLUX-mimic: A video-action model, developed in collaboration with mimic, running robots that have been tested and deployed at Audi.<br>b]:font-medium [&>strong]:font-medium">At first glance, producing convincing visual content and controlling robots seem to have little in common. One requires generating pixels, the other an understanding of how the physical world responds when you touch and manipulate it. If one model does both, it was never really only a content creation model. It is a model of how the world behaves, and content creation is one thing one can do with it.<br>b]:font-medium [&>strong]:font-medium">That is what FLUX 3 is.<br>Video is the hard part<br>b]:font-medium [&>strong]:font-medium">FLUX 3 is one model, jointly trained across images, video and audio from the beginning. The most demanding part of that training - accounting for over 95% of the total compute costs - is video prediction. To generate realistic videos, a model has no choice but to learn contact, motion, weight, cause and effect; get any of them wrong and it looks wrong. Learning to render the world accurately means learning how the world behaves.<br>b]:font-medium [&>strong]:font-medium">Relatively speaking, audio is the easy modality. Low dimensional and far less detailed than video, it makes up less than 0.5% of the tokens in a 720p video with audio. Once a model has done the hard work of learning video understanding, it will learn the causal relationship between video and audio to predict speech synchronized to lip movement and audio effects synchronized to the physical events causing them.<br>b]:font-medium [&>strong]:font-medium">Actions follow the same shape: a low dimensional representation of a robot's state, tightly coupled to visual observations. Actions, audio and video frames are all partial representations of a single underlying physical reality. After the model has learned about the physical processes behind video and audio, action prediction is not a new departure - it is one more view of the reality it already models.<br>A single backbone<br>b]:font-medium [&>strong]:font-medium">If that framing is correct, teaching FLUX 3 to predict actions should not incur lasting costs: we expect a brief phase of disturbance as the model has to learn the structure of the action space and align its internal representation of the world to it, before returning to full performance. That is exactly what we observe.<br>b]:font-medium [&>strong]:font-medium">In a large-scale training run, we added action prediction to the curriculum and observed the effect on video generation quality. Human ratings on text-to-video and image-to-video initially fell by up to 10% as the model started to incorporate the new action modality. After 3500 steps, the model had regained its full previous quality on video generation tasks while now also predicting actions.

b]:font-medium [&>strong]:font-medium">Each series is normalized to its own quality before action prediction was added. Higher is better.<br>b]:font-medium [&>strong]:font-medium">The model had to integrate actions into its inputs and outputs - but doing so didn't cost it capacity permanently. It merely had to learn how this new modality relates to its existing model of the world. Once this was figured out, the performance penalty on its existing capabilities was gone. Video generation and action prediction don't need separate foundations. The same backbone carries both.<br>b]:font-medium [&>strong]:font-medium">This makes Physical AI a natural extension of our roadmap at Black Forest Labs rather than a change in direction. Content creation is what our multimodal FLUX 3 backbone does with image, video and audio. Physical AI is what it does with actions. One foundation model, with visual intelligence at its core, enabling two families of applications. We didn't build a separate foundation model. We focused on the hard thing: building a model that understands the world. Acting in it is what that understanding makes possible.<br>From lab to reality: FLUX-mimic<br>b]:font-medium [&>strong]:font-medium">What happens when we point the FLUX 3 backbone at real automation tasks on real production lines?<br>b]:font-medium [&>strong]:font-medium">That's the question...

font medium video model flux strong

Related Articles