StepZen's StepAudio 3 Voice Models Claim Multiple Global Firsts
StepZen Introduces StepAudio 3 Series, Promising Breakthroughs in Voice AI
On September 15, StepZen launched its new StepAudio 3 series of voice large models, a collection of five products aimed at different audio tasks. The lineup includes StepAudio3Realtime, StepAudio3ASR, StepAudio3TTS, StepAudio3Gen, and StepAudio3Music. Several of these models have already claimed the top spot on the authoritative Artificial Analysis leaderboard, according to the company. All five are now available on the StepZen open platform.
The series targets five core scenarios: realistic voice generation, comprehensive audio content creation, real-time voice interaction, complex voice understanding, and music creation.
Real-Time Conversation Gets a Major Upgrade
StepAudio3Realtime is built for full-duplex, native real-time conversation. It scored 98.9% on the Artificial Analysis Conversational Dynamics ranking, placing first globally, and also topped the Speech Reasoning ranking with 99.7% accuracy. The model can handle interruptions, continuous feedback, and understands not just words but also tone, emotion, paralanguage, and background sounds. It runs reasoning and voice generation in parallel and manages tool calls and long tasks asynchronously, so conversations don't stall.
Speech Recognition That Goes Beyond Transcription
StepAudio3ASR combines high-precision speech recognition with the reasoning power of large language models. It handles Chinese, English, dialects, mixed language, long audio, and specialized fields like medicine, law, finance, automotive, and programming. Even tricky inputs—low volume, fast speech, unclear pronunciation, singing, or background music—are no problem. Its non-streaming word error rate is just 1.7%, tying for first worldwide.
Voice Generation That Sounds Human
StepAudio3TTS is designed for realistic voice generation in real-time interactions. It mimics natural speech patterns, including intonation, rhythm, and pauses. It can reproduce laughter, hesitation, stuttering, repetition, and self-corrections, and even adjusts emotional tone based on context. The streaming architecture allows generation and playback to happen simultaneously.
One-Shot Audio Content Creation
StepAudio3Gen simplifies audio production. Instead of a long chain of recording, editing, and mixing, users just provide a natural language description and reference audio. The model then generates vocals, sound effects, ambient sounds, and background music all at once. It offers fine control over character voice, speaking style, emotion, dialects, and laughter, and can specify the timing and order of dialogue, effects, and music, essentially acting as the sound designer for the entire piece.
Music Creation Made Accessible
StepAudio3Music supports both zero-shot generation and multi-round interactive creation using ABC notation. Users can input lyrics, a cappella, reference songs, or notation, and control genre, vocals, melody, rhythm, instruments, and emotions through natural language. The model deeply understands song structure, cross-paragraph melodic development, and energy changes. In a cappella accompaniment mode, a single vocal track can automatically generate harmony, rhythm, instruments, and a full arrangement—entering real music production territory.
Key Points
- StepZen's StepAudio 3 series includes five models for real-time conversation, speech recognition, voice generation, audio content creation, and music.
- StepAudio3Realtime and StepAudio3ASR both rank first globally on Artificial Analysis leaderboards.
- The models are available now on the StepZen open platform.
- They aim to make voice interaction more natural and audio production more accessible.