Black Forest Labs Unveils Flux3: AI That Generates 20-Second Video with Sound
German AI startup Black Forest Labs (BFL) has officially released Flux3, a multimodal foundation model that pushes the boundaries of AI-generated media. Unlike many models that handle video and audio separately, Flux3 produces synchronized audio-video clips up to 20 seconds long in a single pass.

What Makes Flux3 Different?
Flux3 is built on a novel Self-Flow architecture, which integrates dedicated codecs for images, video, audio, and motion. This allows the model to understand and generate content across physical and digital environments seamlessly. It supports a range of functions: text-to-video, image-to-video, video-to-video, keyframe-based transitions, and even multilingual dialogue.
Outperforming the Competition
In early benchmark tests at 720p resolution with 10-second clips, Flux3 delivered impressive results. It beat Luma's Ray3.2 with a 93% win rate and crushed Runway's Gen-4.5 with a 77% win rate. It also held a slight edge over other top-tier models like Seedance 2.0 and Gemini Omni Flash. These numbers suggest Flux3 is a serious contender in the AI video generation space.
Stepping into the Real World
BFL isn't stopping at digital content. The company partnered with Mimic Robotics to develop Flux-mimic, a video action model designed for robotics. The system has already begun production task testing at an Audi factory, marking a significant step from pure AI generation to real-world manufacturing applications.
Release Strategy
BFL is rolling out Flux3 in phases. The Flux3 Video component is available now, while Flux3 Image and the open-source "Flux3 Dev" weights are expected in the near future. This staggered approach allows developers and researchers to explore the model's capabilities step by step.
Key Points:
- Flux3 generates up to 20 seconds of synchronized audio and video natively.
- Built on Self-Flow architecture with dedicated codecs for multiple modalities.
- Outperforms Luma Ray3.2 (93% win rate) and Runway Gen-4.5 (77% win rate) in early tests.
- Supports text/image/video to video, keyframe transitions, and multilingual dialogue.
- Robotics variant Flux-mimic is being tested in an Audi factory for production tasks.
- Flux3 Video is available now; Image and open-source versions coming soon.