Black Forest Labs has officially launched Flux 3, marking a decisive step beyond image generation. The new model can produce 20-second video clips with synchronized audio from text prompts, and has also been adapted to control robotic arms. In collaboration with Swiss robotics company Mimic Robotics AG, the team created the Flux-mimic model, which is already being tested in Audi factories for complex assembly tasks such as installing flexible door seals.
The essentials
- Flux 3 generates videos up to 20 seconds long with synchronized audio – true multimodal text-to-video
- In blind tests, evaluators preferred Flux 3 in 77 % of comparisons against Runway Gen-4.5 and 93 % against Luma Ray 3.2
- Flux-mimic controls robotic arms with reaction times around 101 milliseconds – already in production at Audi
- Full open-source version arrives in late 2026
From images to video: The technical leap
Flux 3 was trained on image, video, and audio data within a single architecture – genuine multimodality rather than chaining separate tools. The results speak clearly. According to the source, blind tests demonstrate strong competitive performance. The numbers are striking:
| Comparison | Flux 3 Advantage |
|---|---|
| vs. Runway Gen-4.5 | 77 % |
| vs. Luma Ray 3.2 | 93 % |
| vs. Gemini Omni / Seedance | 52 % |
Moreover, Flux 3 retains the strengths of its predecessors in static image generation – Black Forest Labs showcases examples ranging from photorealism to diverse artistic styles.
Robotics as the next frontier
The most compelling aspect of this launch, however, is not video generation alone, but the leap into the physical world. Building on Flux 3's video prediction engine, Black Forest Labs developed the Flux-mimic action model with Mimic Robotics. This can directly control robotic arms – with impressive precision. At Audi, the systems are already being deployed for complex assembly tasks like installing flexible door seals, with reaction times around 101 milliseconds.
CEO and co-founder Robin Rombach captures the underlying philosophy:
"A model that only learns images can only generate images."
The core thesis: learning to predict video means learning the physical principles behind it – weight, contact, timing. That's exactly what machines need to move through the real world.
Comeback with caveats
This launch marks a major comeback for Black Forest Labs. Founded in 2024, the company quickly gained attention with open-source image models that outperformed Stable Diffusion. With Flux 3, the team is now making a bold statement – though with a caveat: a fully open-source version of the model won't be available until late 2026. Until then, Flux 3 remains proprietary.
What this means for enterprises
For German manufacturers and machine builders, Flux 3 could prove valuable. The system demonstrates that AI models are no longer just creative tools – they can work directly in production with low latency and high precision. The Audi tests in particular suggest such systems are production-ready. However, questions remain about availability and cost for other manufacturers. And for those betting on open-source: the wait until 2026 is significant.
Sources
Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.




