Skip to main content

Black Forest Labs unveils Flux3: AI that sees and hears at the same time

German AI startup Black Forest Labs just dropped something that might make you do a double-take. Meet Flux3, a multimodal foundation model that doesn't just generate images or videos—it produces synchronized audio and video clips up to 20 seconds long. Think of it as an AI that can both see and hear, and keep them in perfect sync.

What makes Flux3 different?

Most AI models today are specialists. You've got your image generators, your video makers, your audio creators. Flux3 is a generalist. Built on something called the Self-Flow architecture, it packs dedicated codecs for images, video, audio, and motion. The result? A single model that understands and generates content across multiple senses.

The headline feature is audio-video synchronization. Flux3 is the first native multimodal model to generate audio that matches the visuals—no post-production lip-syncing or manual alignment needed. It can handle text-to-video, image-to-video, video-to-video, keyframe transitions, and even multi-character dialogue.

Image

How does it stack up?

Early benchmarks are impressive. In tests with 720p resolution and 10-second clips, Flux3 beat Luma Ray3.2 with a 93% win rate. It also outperformed Runway Gen-4.5 by 77%. Even against heavyweights like Seedance2.0 and Gemini Omni Flash, it held a slight edge.

From screen to factory floor

Flux3 isn't just about entertainment. Black Forest Labs teamed up with Mimic Robotics to create Flux-mimic, a version tailored for robotics. It's already running production tasks at an Audi factory. That's right—multimodal AI has moved from generating videos to guiding real-world machines.

What's next?

Black Forest Labs is rolling out Flux3 in stages. Flux3Video is already available. Flux3Image and an open-source version called Flux3Dev are coming soon. If you're curious about where AI is headed, this is a glimpse: models that don't just generate content, but understand the world the way we do—through multiple senses at once.

Key Points

  • Flux3 generates up to 20 seconds of synchronized audio-video content natively.
  • It uses a Self-Flow architecture with dedicated codecs for different modalities.
  • Outperforms Luma Ray3.2 (93% win rate) and Runway Gen-4.5 (77% advantage) in early tests.
  • A robotics variant, Flux-mimic, is already deployed in an Audi factory.
  • Flux3Video is live; Flux3Image and open-source Flux3Dev are on the way.