Skip to main content

MiniMax Music3: Turn Lyrics into Full Songs Up to 5 Minutes Long

MiniMax Music3: Turn Lyrics into Full Songs Up to 5 Minutes Long

Imagine typing in a few lines of lyrics, describing the vibe you want—maybe a dreamy synth-pop ballad or an upbeat indie rock track—and then hitting "generate." Within moments, you get a complete song, up to five minutes long, with vocals, instrumentation, and a full arrangement. That's exactly what MiniMax's new Music3 model promises.

How It Works: A Tale of Two Models

At the heart of Music3 is a clever division of labor. Instead of one massive model trying to do everything, MiniMax uses two specialized models that work together like a composer and an orchestrator.

The first, called the Global LLM, is the big one—8 billion parameters. It's responsible for the song's overall structure and long-term flow. Think of it as the architect who decides where the verse goes, when the chorus hits, and how the bridge builds tension. It's initialized from Qwen3-8B, a powerful language model, and adapted to understand musical semantics.

The second, the Local LLM, is much smaller—just 0.6 billion parameters. But don't let its size fool you. It handles the fine-grained acoustic details, filling in the nuances that make a song feel alive: the subtle vibrato in a vocal, the shimmer of a cymbal, the warmth of a bassline. Together, they create a track that's both structurally sound and sonically rich.

What You Get: A Studio-Quality WAV

The output is a 32kHz, 16-bit stereo WAV file—essentially CD-quality audio. That's a big deal for musicians and content creators who want to use AI-generated music in their projects without worrying about low-fidelity artifacts.

But the real magic is in the model's ability to maintain coherence over long durations. MiniMax claims that Music3 can hold onto the musical theme, rhythm, vocal identity, and arrangement progression throughout the entire track. No more songs that start strong but fall apart halfway through. The model handles all the standard sections—intro, verse, pre-chorus, chorus, bridge, instrumental break, and outro—without missing a beat.

Why This Matters

For independent artists, this could be a game-changer. Stuck on a chord progression? Need a demo to pitch to a producer? Music3 can generate a full arrangement in seconds, giving you a solid foundation to build on. For content creators, it means royalty-free music that's tailored to your specific needs, whether it's a podcast intro or a background score for a video.

Of course, there are questions. Will AI-generated music ever replace human creativity? Probably not. But it can certainly augment it, offering new tools for exploration and inspiration. And as the technology improves, the line between human and machine-made music will only blur further.

Getting Started

MiniMax has already released the model page on GitHub, complete with sample songs you can listen to. So if you're curious, head over and give it a try. Who knows? Your next favorite song might be one you wrote with a little help from AI.


Key Points

  • MiniMax Music3 generates complete songs up to 5 minutes long from lyrics and a description.
  • Uses a hierarchical autoregressive architecture with two models: an 8B-parameter Global LLM for structure and a 0.6B-parameter Local LLM for acoustic details.
  • Outputs 32kHz, 16-bit stereo WAV files with high fidelity.
  • Maintains musical theme, rhythm, vocal identity, and arrangement throughout the track.
  • Available on GitHub with sample songs for listening.

Image