Skip to main content

Alibaba's Qwen3.8-Omni-Flash Cuts Audio Costs by 98%

Alibaba's Qwen3.8-Omni-Flash: A Multimodal Powerhouse with a Price Tag That Turns Heads

Alibaba's Qwen team has quietly released a new native multimodal model, Qwen3.8-Omni-Flash, and it's making waves—not just for what it can do, but for how little it costs to do it.

One Model, Four Inputs

Forget stitching together separate speech and image recognizers. Qwen3.8-Omni-Flash processes text, images, audio, and video natively, all within a single model. It comes with a 1 million token context window—enough to swallow entire meetings or lengthy video archives without constant slicing and dicing. Compared to its predecessor, Qwen3.5-Omni-Plus, Alibaba claims an average improvement of over 25% across 29 benchmark tests.

Image

Audio: Where It Shines

The model's audio chops are particularly impressive. In the WildClawBench-MM multimodal tool calling test, it scored 71.0, beating Gemini3.8Flash's 58.9. On DailyOmni, it edged out Gemini 85.1 to 84.0; on SpotSoundBench, it dominated 67.2 to 39.7; and on MMAU, it won 81.8 to 76.9. The jump in multi-speaker meeting recognition is even more dramatic: speaker error rate (DER) plummeted from 88.11 in the previous generation to 3.35, and word error rate with speaker attribution (cpWER) fell from 89.61 to 17.18.

But it's not a clean sweep. On AgenticVBench, Gemini scored 45.0 versus Qwen's 36.8; on OmniGAIA, Gemini led 78.6 to 74.0. And in some video understanding tasks, Gemini still holds the edge. So while the "approaching Gemini" narrative holds for audio and audio-video capabilities, it's not a universal victory.

The Price Shock

Here's the headline that might actually change how developers work: audio input costs have dropped by more than 98% per hour, and combined audio-video input is down over 93%. Alibaba Cloud's international region pricing is $0.15 per million input tokens, $0.016 per million cached input tokens, and $0.47 per million output tokens. If those numbers hold up, analyzing long meetings, podcasts, live streams, and massive video libraries becomes almost trivially cheap.

Built for Agents

Qwen3.8-Omni-Flash isn't just a passive listener. It can plan tasks, call tools, and complete work. The team is targeting "agent delivery": the model can analyze long videos, pinpoint key moments, then use tools to edit video, organize content, generate meeting summaries, and create subtitles. It handles audio in 113 languages and dialects, supports stereo and 4-channel spatial audio analysis, and offers function calls, online search, and context caching.

Accompanying tools include Qwen-MM-Plugins for multimodal agent development, with Qwen-Live-Harness on the horizon. The model is available through Alibaba Cloud BaiLian in regions including Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, and Virginia. The model itself isn't open-sourced, but the release opens up the accompanying multimodal plugins and development tools.

Key Points

  • Native multimodal: Handles text, images, audio, and video in one model with a 1M token context window.
  • Audio leadership: Outperforms Gemini3.8Flash on four key audio benchmarks, with massive gains in multi-speaker recognition.
  • Not universal: Gemini still leads in some video and agentic benchmarks.
  • Price disruption: Audio input costs cut by over 98%; audio-video combined by over 93%.
  • Agent-ready: Supports task planning, tool calling, and 113 languages for audio input.
  • Availability: Accessible via Alibaba Cloud BaiLian in multiple global regions; plugins and dev tools open-sourced.

If the cost reduction holds, the economics of automatically analyzing long meetings, podcasts, and video archives change overnight. And the rivalry between Qwen and Google Gemini in the multimodal arena just got a lot more interesting.