Skip to main content

Black Forest Lab's FLUX3: One Model, 20-Second Audio and Video, Outshining Grok and Seedance

Black Forest Labs has just dropped FLUX3, a new multimodal foundation model that's already making waves. Available in early access, FLUX3 can generate up to 20 seconds of video along with synchronized audio in a single pass. That's a big deal—most models struggle to produce even a few seconds of coherent video, let alone with sound that matches.

The secret sauce? A unified architecture that learns from images, videos, and audio simultaneously. Built on a self-supervised flow matching framework called Self-Flow, FLUX3 extends the company's earlier FLUX.1 and FLUX.2 series into full multimodal territory. Think of it as a model that doesn't just see or hear—it does both at once, and it's surprisingly good at it.

In head-to-head tests, FLUX3 beat Grok Imagine Video with a 69% win rate on 10-second 720p clips with sound. It also edged out Seedance 2.0 and Gemini Omni Flash, scoring a 52% win rate. Those numbers suggest FLUX3 isn't just keeping up—it's setting a new bar for quality.

But it's not just about video. FLUX3 supports a wide range of creation modes: text-to-video, image-to-video, video-to-video, and even keyframe-to-video. You can also extend existing videos or audio, create multilingual dialogues, and stitch multiple shots together. The model handles all of this natively, without needing separate tools or post-processing.

And here's a twist: FLUX3 is also being used for robot behavior prediction. That's right—the same model that generates lifelike videos can help robots understand and anticipate actions. It's a surprising crossover, but it shows how versatile the underlying technology is.

Image

Of course, early access means we're still in the testing phase. But if the benchmarks hold up, FLUX3 could change how we think about AI-generated media. No more awkward silent clips or mismatched audio—just one model, one generation, and a whole lot of potential.

Key Points

  • FLUX3 generates up to 20 seconds of video with native audio in a single pass.
  • Uses a unified architecture and self-supervised flow matching for joint learning of images, video, and audio.
  • Outperforms Grok Imagine Video (69% win rate) and Seedance 2.0/Gemini Omni Flash (52% win rate) in quality tests.
  • Supports multiple creation modes: text-to-video, image-to-video, video-to-video, keyframe-to-video, and more.
  • Also applied to robot behavior prediction, showcasing its versatility.