Skip to main content

Black Forest Labs Unveils Flux3: AI That Sees and Hears Together

German AI startup Black Forest Labs has officially launched Flux3, a multimodal foundation model that marks a significant leap in AI's ability to understand and generate content across different media. Unlike previous models that specialized in either images or videos, Flux3 integrates native audio generation, allowing it to produce synchronized audio-video clips up to 20 seconds long.

Built on the company's Self-Flow architecture, Flux3 uses dedicated codecs for images, video, audio, and motion. This unified approach means the model can handle text, image, and video-to-video conversion, keyframe transitions, and even multi-character dialogue—all within a single framework.

Image

In early benchmarks, Flux3 has shown impressive performance. At 720p resolution and 10-second clip settings, it defeated Luma Ray3.2 with a 93% win rate and outperformed Runway Gen-4.5 by 77%. Even against top-tier models like Seedance2.0 and Gemini Omni Flash, Flux3 maintained a slight edge.

But Black Forest Labs isn't stopping at media generation. In collaboration with Mimic Robotics, the company developed Flux-mimic, an action model tailored for robotics. This version is already being tested on production tasks at an Audi factory, bringing multimodal AI from the screen to the factory floor.

The release is being rolled out in stages: Flux3Video is available now, with Flux3Image and an open-source version called "Flux3Dev" coming soon. This phased approach allows developers and researchers to start experimenting with the video capabilities while waiting for the full suite.

Key Points

  • Flux3 is the first multimodal foundation model with native audio generation, capable of producing up to 20 seconds of synchronized audio-video content.
  • It uses a unified Self-Flow architecture with dedicated codecs for different media types.
  • In benchmarks, Flux3 outperformed Luma Ray3.2, Runway Gen-4.5, and held its own against Seedance2.0 and Gemini Omni Flash.
  • A robotics-focused version, Flux-mimic, is already in use at an Audi factory.
  • The model is being released in stages: Flux3Video now, with Flux3Image and open-source weights to follow.