Skip to main content

Alibaba's Qwen Drops Five Speech Models, Slashes TTS and ASR Prices

Alibaba's Qwen Drops Five Speech Models, Slashes TTS and ASR Prices

On September 23, Alibaba's Qwen team unveiled a major upgrade to its audio toolkit, launching five new speech models under the Qwen-Audio-3.1 banner. These models span the entire audio spectrum—from recognition and synthesis to real-time interaction and creative audio generation. And the pricing? It's a game-changer: text-to-speech (TTS) costs have been cut by about 70%, real-time services by 85%, and automatic speech recognition (ASR) by a staggering 95%. This move makes advanced voice capabilities accessible to developers at an unprecedented low cost.

Speech Recognition That Understands Context

At the heart of the update is Qwen-Audio-3.1-ASR, a next-gen speech recognition model that doesn't just transcribe—it comprehends. Its standout feature is native transcription polishing, which automatically removes filler words, eliminates repetitions, and restructures sentences for smoother, more logical output. Role-based transcription accurately identifies who said what and when, making it perfect for meetings, interviews, and dialogue analysis. The system outputs speaker labels, timestamps, and text, handling interruptions and overlapping speech with ease.

But it's not just about English. The model supports 30 languages and 16 Chinese dialects, recognizes industry jargon and hot words, and maintains consistent terminology across long audio contexts. With a first-word response delay of just 160 milliseconds, real-time transcription feels seamless. Performance metrics are impressive: on public dialect datasets, the average word error rate is 4.55%, and for 16 Chinese dialects, it's 10.38% in internal tests. Even challenging dialects like Wenzhou see significant improvements, with an 82.10% semantic accuracy rate when converting dialects to Mandarin.

Image

Taking it a step further, Qwen-Audio-3.1-ASR-Next expands recognition from mere speech to full audio understanding. It can detect emotions, environmental sounds, and mechanical noises, perform sound descriptions, event localization, and even answer questions about audio content. Imagine a clip with dialogue, background music, and ambient noise—it not only provides subtitles but also tells you what sounds are present and when events occur. In role-based ASR tests, ASR-Flash-Next and ASR-Flash scored top marks in five and three out of eight metrics respectively, surpassing previous benchmarks.

Speech Synthesis with Realistic Expression

For speech synthesis, Qwen-Audio-3.1-TTS focuses on natural expression, not just pronunciation accuracy. It supports multilingual and dialectal synthesis, enabling voice migration across languages so users can speak multiple languages instantly. You can control emotion, speech rate, and expression through instructions, tailoring the voice to specific content and scenarios.

The real surprise is Qwen-Audio-3.1-TTS-Next, which elevates audio creation from single-voice synthesis to a general audio generation system. Previously, creating a complete audio piece required a team—hosts, sound engineers, mixers, editors. Now, one model can generate human voices, sound effects, and ambient sounds based on text, timestamps, and reference audio. It supports multiple voice replication and multi-turn dialogues, keeping each character's voice consistent. It can merge various sounds, adding layers of material, distance, and space—imagine a podcast with narration, two-character dialogue, and background city hums, evening winds, and can-crushing sounds, making the conversation feel like it's happening in a real scene. It also accepts natural language control for tone, speed, volume, music style, and instrumentation, with fine-grained timestamps and 48kHz output. This precision meets the needs of podcasts, audiobooks, films, games, and ads.

Image

Real-Time Interaction That Feels Human

Finally, Qwen-Audio-3.1-Realtime breaks the chain of separate recognition, modeling, and synthesis. It operates in full-duplex mode, allowing simultaneous speaking and listening, so users can interrupt or interject anytime—just like a real conversation. The model understands not only words but also the emotions and intentions behind sounds. Sense a depressed tone? It won't mechanically reply with "I understand," but instead slow down and adjust its wording. It switches languages without losing context, tone, or rhythm. It avoids fabricating uncertain information and respects boundaries for high-risk content. Using a multi-teacher distillation architecture, it improves reasoning and safety while maintaining low latency. Plus, it can call Agent tools during conversations, accessing APIs, knowledge bases, and business systems to retrieve and execute results, seamlessly integrating them back into the dialogue.

Key Points

  • Five new models under Qwen-Audio-3.1 cover recognition, synthesis, real-time interaction, and audio creation.
  • Price cuts: TTS down 70%, real-time down 85%, ASR down 95%.
  • ASR supports 30 languages, 16 Chinese dialects, with context-aware transcription and role labeling.
  • TTS offers multilingual synthesis, voice migration, and emotion control; TTS-Next enables full audio scene generation.
  • Realtime model features full-duplex conversation, emotional understanding, and Agent tool integration.