Skip to main content

Alibaba's Qwen Drops Five New Speech Models—ASR Price Plummets 95%

Alibaba's Qwen team has just unveiled a suite of five new speech models under the Qwen-Audio-3.1 banner. This isn't just a minor update—it's a full-on upgrade that touches every corner of audio AI, from listening and understanding to generating and interacting. And to sweeten the deal, they've slashed prices across the board: text-to-speech (TTS) is down about 70%, real-time interaction down 85%, and automatic speech recognition (ASR) takes a dramatic 95% cut.

Five Models, One Big Leap

The star of the show might be Qwen-Audio-3.1-ASR, a next-gen speech recognition model that does more than just transcribe. It automatically polishes your transcripts—removing filler words like "um" and "uh," tidying up repetitions, and even reorganizing sentences for clarity. Need to know who said what in a meeting? Its role-based transcription captures speaker turns with timestamps, giving you a structured, searchable record. The model handles 30 languages and 16 Chinese dialects, plus industry jargon and hotwords. Latency? Just 160 milliseconds for the first character. In tests, it achieved an average character error rate (CER) of 4.55% on 11 subsets of KeSpeech and WSYue, and 10.38% on internal dialect sets. Semantic sentence accuracy on AST hit 82.10%.

But wait, there's more. ASR-Next, built on a new architecture, goes beyond words to understand emotions, background noises, and even mechanical sounds. It can describe audio scenes, pinpoint events, and answer questions about what it hears. In role-based ASR evaluations, the Flash-Next and Flash versions notched five and three top scores respectively, beating the previous Fun-ASR system.

On the synthesis side, Qwen-Audio-3.1-TTS lets you generate speech in multiple languages and dialects, with the same voice smoothly switching between them. You can tweak emotion, speed, and expression via simple commands. The new TTS-Next takes it further: by unifying text, timestamps, and reference audio, it can produce human voices, sound effects, and ambient sounds all at once. It supports 48kHz output and fine-grained timestamp control—perfect for podcasts, audiobooks, films, and games.

Image

From Understanding to Creation: Voice as the New AI Interface

Then there's Qwen-Audio-3.1-Realtime, a full-duplex model that lets you speak and listen simultaneously. It handles interruptions gracefully and goes beyond mere text recognition—it picks up on your emotions and intent. If it senses you're feeling down, it might slow down and choose gentler words. It also switches languages on the fly and can call APIs, knowledge bases, or business tools mid-conversation. Built on a multi-teacher distillation architecture, it keeps latency low while boosting factual reliability and safety.

Image

All these models form a complete audio stack: understanding, generation, interaction, and creation. With prices slashed, the barrier to entry is lower than ever. Whether you're building a voice assistant, a podcast tool, or a real-time translation service, Qwen's latest offerings are worth a close look.

Key Points:

  • Five new Qwen-Audio-3.1 models cover ASR, TTS, real-time interaction, and audio creation.
  • ASR price drops by 95%, TTS by 70%, and Realtime by 85%.
  • ASR supports 30 languages, 16 Chinese dialects, and offers role-based transcription with 160ms latency.
  • TTS-Next generates voices, sound effects, and ambient audio in one go, with 48kHz output.
  • Realtime model is full-duplex, emotion-aware, and can call external tools during conversation.