Skip to main content

Qwen3.8-LiveTranslate Cuts Real-Time Translation Lag to 2.3 Seconds

Qwen3.8-LiveTranslate Cuts Real-Time Translation Lag to 2.3 Seconds

Imagine watching a live international broadcast where the speaker's words flow into your ears in your language, almost instantly—and in their own voice. That's the promise of Alibaba's latest real-time translation model, Qwen3.8-LiveTranslate, which just got a major upgrade.

The new version tackles four key areas: translation quality, latency, speaker identification, and speech synthesis. Under the hood, it uses an audio-text interleaved architecture that merges streaming understanding, text output, and speech generation into a single causal sequence. The result? Average latency drops from 2.8 seconds in the previous generation to 2.3 seconds—a noticeable improvement when you're trying to follow a fast-paced conversation.

Image

Three New Tricks That Make It Feel Human

First, real-time speaker separation clearly identifies who's talking during multi-person exchanges. Even better, the translated speech retains the original speaker's voice timbre, so you hear the same person, just in a different language. Second, the model simultaneously displays both source and target text—handy if you want to double-check a phrase or learn as you listen. Third, long-context disambiguation uses cross-turn context to resolve ambiguities from proper nouns and pronouns, cutting down on those awkward mistranslations.

Image

Under the Hood: Thinker and Talker

The model relies on a Hybrid MoE Thinker–Talker dual-module design. The Thinker handles understanding and translation, while the Talker generates the translated speech with the original voice timbre. On the Omnilingua-MSpeaker multi-speaker long audio evaluation set—covering 14 language directions—the model outperforms current mainstream real-time translation systems in translation fidelity, fluency, conciseness, and speaker separation error rate.

Right now, Qwen3.8-LiveTranslate supports 60 languages, and the API is already available on the Qwen AI platform. The team says the next steps will focus on further reducing end-to-end latency, adding cross-conversation long-term memory, and expanding language support.

So, what does this mean for everyday users? If you've ever struggled with clunky translation tools that sound like robots, this upgrade brings us closer to seamless, natural communication across languages. It's not perfect yet, but it's a big step toward breaking down language barriers in real time.

Key Points

  • Latency reduced to 2.3 seconds from 2.8 seconds in the previous generation.
  • Real-time speaker separation identifies who's speaking and preserves their voice timbre in translation.
  • Simultaneous source and target text display for easy verification.
  • Long-context disambiguation reduces errors from proper nouns and pronouns.
  • Hybrid MoE Thinker–Talker design powers understanding, translation, and speech synthesis.
  • Supports 60 languages; API available now on the Qwen AI platform.
  • Next up: lower latency, cross-conversation memory, and more languages.