Skip to main content

Alibaba's New TTS Model Speaks 20 Dialects, First Word in 300ms

Alibaba's Qwen team has officially launched its latest real-time speech synthesis model, Qwen-Audio-3.0-TTS, pushing the boundaries of what AI voices can do. The model comes in two flavors: a Flash version optimized for real-time interaction with an initial latency of about 300ms—perfect for smart assistants and other low-latency applications—and a Plus version that prioritizes high-quality generation, offering superior naturalness and voice similarity.

Image

The Plus version has already made waves by ranking first globally in the Speech Arena, a prestigious benchmark by Artificial Analysis, outperforming mainstream models like Gemini 3.1 TTS and ElevenLabs v3. But what really sets this model apart are its four core breakthroughs.

Multilingual and Dialect Coverage

First up, the model supports 16 languages, with the Plus version achieving an average speaker similarity of 82.75% across all of them—industry-leading performance. But here's the kicker: it also supports 20 Chinese dialects, from Cantonese to Shanghainese and beyond. Unlike many models that water down dialect characteristics, this one keeps them authentic, making it sound like a native speaker.

Natural Language Instructions

Gone are the days of needing technical jargon to control voice output. With Qwen-Audio-3.0-TTS, you can use free-style natural language instructions like "gentle customer service tone" or "live-streaming host style" to generate the desired speech. No professional annotations required—just tell it what you want, and it delivers.

Fine-Grained Tag Control

For those who need precise control, the model supports structured tags such as [gasp] and [angry]. This allows you to inject non-verbal details like breathing, laughter, or emotional cues into the speech, making it ideal for gaming, audiobooks, and other immersive experiences.

Acoustic Robustness

Even if your reference audio is noisy or reverberant, the model can automatically filter out background noise while preserving voice quality. This ensures stable synthesis results regardless of the input conditions—a huge plus for real-world applications.

Image

The accompanying premium voice library covers a wide range of voice types, including instructions, dialects, and small languages. It supports 48K high-definition audio output (expected to be available on July 24th) and can synthesize up to 3 minutes of long text in a single session. The model is now fully accessible on the Alibaba Cloud BaiLian platform, inviting developers to explore new possibilities in voice interaction.

Key Points

  • Two versions: Flash (300ms latency) for real-time use, Plus for high-quality generation.
  • Global #1: Plus version tops Speech Arena, beating Gemini 3.1 TTS and ElevenLabs v3.
  • Dialect support: 20 Chinese dialects with authentic characteristics.
  • Natural control: Use everyday language to command voice styles.
  • Robustness: Handles noisy audio gracefully.
  • Availability: Now on Alibaba Cloud BaiLian platform.