Skip to main content

Qwen-Audio-3.1 Rolls Out Five Speech Models with Price Cuts Up to 95%

Qwen-Audio-3.1 Launches: A Full Audio Stack, Now Cheaper

On September 23, Alibaba's Qwen large model team unveiled the Qwen-Audio-3.1 series—five audio models that together cover everything from listening to speaking, creating, and even chatting in real time. The lineup includes speech recognition (ASR), audio understanding (ASR-Next), text-to-speech (TTS), audio creation (TTS-Next), and real-time interaction (Realtime). Most are already available via API on the Qwen AI platform, with ASR-Next coming soon.

But the real headline? The price cuts. TTS is down about 70%, Realtime by 85%, and ASR by as much as 95%. That's not a typo—ASR is practically free compared to before.

Image

What Each Model Brings to the Table

ASR now handles 30 languages and 16 Chinese dialects, with a first-word response time of just 160 milliseconds for streaming. It can automatically clean up transcripts—removing filler words and repetitions—and even tag speakers with timestamps. Think of it as a transcriptionist that never gets tired.

ASR-Next goes beyond words. It picks up emotions, background sounds, and mechanical noises, and can answer questions about audio content. Imagine asking your recording, "What was that crash?" and getting an answer.

TTS lets you clone voices across languages and control emotion and speed with simple instructions. TTS-Next is for creators: it generates human voices, sound effects, and ambience together, supports multi-role replication, and outputs at 48kHz. Perfect for podcasts, audiobooks, or game audio.

Realtime enables full-duplex conversation—you can interrupt it anytime, switch languages mid-sentence, and even trigger tool calls. It's already integrated with agents like Qoder and Qwen Office, plus smart hardware like Qwen AI glasses.

Performance That Speaks for Itself

In tests, ASR hit an average character error rate (CER) of 4.55% on open-source dialect sets and 10.38% on self-built Chinese dialect sets. For converting dialects to Mandarin, semantic sentence accuracy reached 82.10%. The ASR-Flash-Next and ASR-Flash models ranked first in five and three out of eight metrics across four test sets, beating the previous Fun-ASR cascading system.

The Bigger Picture

Voice is becoming a natural interface for human-computer interaction. With these models, Qwen is betting that talking to machines will feel as easy as talking to a person. And with prices slashed, that future just got a lot more accessible.

Key Points

  • Five models launched: ASR, ASR-Next, TTS, TTS-Next, Realtime.
  • Price cuts: TTS ~70%, Realtime ~85%, ASR up to 95%.
  • ASR supports 30 languages, 16 dialects, 160ms streaming response.
  • ASR-Next understands emotions, ambient sounds, and answers audio questions.
  • TTS-Next creates voices, sound effects, and ambience for podcasts, audiobooks, games.
  • Realtime allows interruptions, language switching, and tool calls.
  • Performance: ASR achieves 4.55% CER on open-source dialects, 82.10% semantic accuracy for dialect-to-Mandarin.
  • Integration: Realtime works with Qoder, Qwen Office, and Qwen AI glasses.