Skip to main content

StepAudio3: Five New Voice Models That Top Global Rankings

StepAudio3 Series: Five Models, One Giant Leap for Voice AI

StepZenith has unveiled its latest innovation: the StepAudio3 series, a suite of five audio models that are already making waves in the AI community. Released on the StepZenith Open Platform, these models cover everything from human-like voice generation to real-time interaction, and they're not just impressive—they're record-breaking.

Meet the Models

StepAudio3Realtime leads the pack with native full-duplex real-time conversation. It doesn't just hear words; it grasps semantics, tone, emotion, even background sounds. It can reason and generate speech simultaneously, handle tool calls, and juggle long tasks asynchronously. In the Artificial Analysis Conversational Dynamics ranking, it scored a stunning 98.9%, claiming the top spot globally. Its voice reasoning accuracy? A near-perfect 99.7%, also number one.

Then there's StepAudio3ASR, which blends speech recognition with large model context understanding. It handles Chinese, English, dialects, mixed languages, long audio, and specialized fields like medical, legal, and financial—all with a word error rate of just 1.7%, tying for first worldwide.

StepAudio3TTS takes voice synthesis up a notch. It captures vocal tone, intonation, rhythm, pauses, and even paralanguage like laughter, hesitation, and stuttering. Plus, it supports streaming generation, making it ideal for real-time interactions.

But wait, there's more. StepAudio3Gen can generate character voices, sound effects, ambient sounds, and background music—all in one go. It even supports multi-character dialogues and voice timing control. And for the musicians out there, StepAudio3Music lets you create songs, a cappella, covers, and multi-round compositions using ABC notation. You can control style, melody, rhythm, instruments, and emotion with natural language.

Image

Why It Matters

The StepAudio3 series isn't just about incremental improvements. It's a leap from single-function voice models to complete audio content production and real-time intelligent interaction. Whether you're building a virtual assistant, creating immersive audio experiences, or composing music, these models are designed to handle it all.

StepZenith's move signals a new era in voice AI—one where machines don't just recognize or generate speech, but understand and interact with the world in a way that feels remarkably human.

Image

Key Points

  • Five models released: StepAudio3Realtime, StepAudio3ASR, StepAudio3TTS, StepAudio3Gen, and StepAudio3Music.
  • Top global rankings: Realtime and ASR models rank first in Artificial Analysis for conversational dynamics and voice reasoning.
  • Versatile capabilities: From real-time conversation to music creation, the series covers a wide range of audio tasks.
  • High accuracy: ASR achieves a 1.7% WER, while Realtime hits 99.7% voice reasoning accuracy.
  • Available now: All models are on the StepZenith Open Platform.