Skip to main content

Black Forest Lab's FLUX3: Generate 20-Second Audio and Video in One Go

Black Forest Labs has just dropped FLUX3, a new multimodal foundation model now available in Early Access. Unlike previous models that specialized in one thing, FLUX3 learns images, videos, and audio together using a unified architecture. It builds on the Self-Flow learning framework—think of it as a self-supervised flow matching system that trains on video, images, and audio simultaneously. This is a big step up from the FLUX.1 and FLUX.2 series, pushing into full-on multimodal generation and understanding.

Image

Native audio, no lip-sync headaches

One of the standout features? FLUX3 can output videos up to 20 seconds long in a single generation, complete with native audio that's synced from the get-go. No more fiddling with separate audio tracks or worrying about mismatched lip movements. It supports a bunch of creation modes: text-to-video, image-to-video, video-to-video, even keyframe-to-video and multilingual dialogue. You can also feed it an existing video and audio to continue the scene, or stitch multiple shots together.

In manual evaluations of 10-second 720p videos with sound, FLUX3 scored a 69% win rate against Grok Imagine Video, and 52% against Seedance 2.0 and Gemini Omni Flash. That's a solid lead across the board.

Beyond video: image smarts and robot predictions

But FLUX3 isn't just about video. Its image capabilities are comprehensive too, and the model is even being applied to robot behavior prediction—a sign that Black Forest Labs is thinking beyond entertainment. The same unified architecture that handles video and audio can also understand and predict physical actions, which could be huge for robotics and autonomous systems.

Key Points

  • FLUX3 generates up to 20 seconds of video with native audio in one go.
  • It uses a unified architecture for joint learning of images, videos, and audio.
  • Outperforms Grok Imagine Video (69% win rate) and Seedance 2.0 (52% win rate) in manual evaluations.
  • Supports multiple creation modes: text-to-video, image-to-video, video-to-video, and more.
  • Extends to robot behavior prediction, showing broader potential.