Skip to main content

Black Forest Labs' Flux3: AI That Sees, Hears, and Syncs

German AI startup Black Forest Labs (BFL) has dropped a new model that's turning heads—and ears. Meet Flux3, a multimodal foundation model that doesn't just generate images or videos; it produces synchronized audio-visual clips up to 20 seconds long. Think of it as an AI that can "see" and "hear" at the same time, blending both into one seamless output.

What Makes Flux3 Different?

Flux3 is built on what BFL calls the Self-Flow architecture. Under the hood, it uses dedicated codecs for images, videos, audio, and motion. This isn't just a patchwork of separate models stitched together—it's a single system that understands and generates multiple modalities at once. The result? Audio that matches the visuals, whether it's a character speaking, a door creaking, or wind rustling through trees.

In early benchmarks, Flux3 flexed its muscles. At 720p resolution and 10-second clips, it beat Luma Ray3.2 with a 93% win rate and outperformed Runway Gen-4.5 by 77%. Even against heavyweights like Seedance 2.0 and Gemini Omni Flash, it held a slight edge. Not bad for a newcomer.

From Screen to Factory Floor

But BFL isn't stopping at entertainment. In a surprising move, the company partnered with Mimic Robotics to develop Flux-mimic, an action model for robotics. This isn't just a lab experiment—it's already running production tasks at an Audi factory. Multimodal AI is moving from your phone screen to the assembly line, and it's happening faster than many expected.

Rollout Plan

BFL is releasing Flux3 in stages. The video component, Flux3Video, is already available. Flux3Image and an open-source version called "Flux3Dev" are coming soon. This phased approach gives developers and creators time to explore each piece before the full suite lands.

Key Points

  • Flux3 is the first native multimodal model to generate synchronized audio-video clips up to 20 seconds long.
  • It uses Self-Flow architecture with dedicated codecs for image, video, audio, and motion.
  • In benchmarks, it outperformed Luma Ray3.2 (93% win rate) and Runway Gen-4.5 (77% win rate).
  • BFL partnered with Mimic Robotics to create Flux-mimic, an action model already deployed in an Audi factory.
  • Release is phased: Flux3Video is live; Flux3Image and open-source Flux3Dev are upcoming.

Image