Qwen's New Omni Model Handles 1M Context, Scores 26% Higher
Qwen3.8-Omni-Flash Arrives with 1M Context and Big Performance Gains
On September 18, Qwen unveiled its latest native multimodal model, Qwen3.8-Omni-Flash, and it's already available for testing on the Qwen AI platform. This model doesn't just handle text—it takes in images, audio, and video too, all while supporting a massive 1 million token context window. That's a lot of room to work with.
But what does it actually do? Beyond the usual programming and GUI tasks, Qwen3.8-Omni-Flash zeroes in on audio and video workflows. Think video editing, MV creation, film production and commentary, audio/video-to-text summaries, and even conversations about audio/video content. These are tasks that demand serious multimodal processing, and Qwen is positioning this model as the go-to for them.
Benchmarks: A 26% Average Leap
Across 30 evaluations, the model scored an average of over 26% higher than its predecessor, Qwen3.5-Omni-Plus. In specific areas like Audio-Visual Agent, Coding, and long-term tasks, the improvements are even more striking:
- WildClawBench-MM jumped by 36.5 points
- AgenticVBench climbed 22.3 points
- UniClawBench hit 69.6 points
On basic capabilities, LongAudioSpan improved by 8.3 points, OmniVideoBench by 9.6 points, and OmniCap-IF saw its CSR/ISR rise by 8.5 and 14.1 points respectively. Perhaps most impressive: the DER/cpWER on AliMeeting dropped from a dismal 88.11/89.61 to just 3.35/17.18. That's not a typo—it's a massive leap in accuracy.

Pricing That Turns Heads
Qwen claims the model's audio and video capabilities are close to Gemini3.8Flash, with overall audio performance actually exceeding it. And the pricing? It's aggressive. The cost per hour for audio input via API has dropped by more than 98%, while audio and video input prices fell by over 93%. For developers and businesses, that's a game-changer.

Built for Long Workflows and Real-Time Interaction
To support extended workflows and real-time interaction, Qwen expanded Qwen-MM-Plugins and open-sourced Qwen-Live Harness. Under the Agentic Understanding mode, OmniVideoBench accuracy rose from 63.4 to 67.8, while token consumption dropped from 145,736 to 79,117—a reduction of about 45.7%. Less compute, better results.

The Model That Helped Build Itself
In a fascinating twist, the team used the model in its own R&D process. They completed evaluation set selection, data construction, and four iterations within 12 hours, building 3,413 training data entries. This self-improvement loop helped reduce the character error rate for Sichuan dialect recognition in Qwen2.5-Omni-3B from 25.79% to 15.30%—a relative decrease of about 40.7%.
Also released alongside is a Realtime version, which Qwen says is the first full-modal large model supporting "sound localization." That means it can pinpoint where a sound is coming from, not just what it is.
Key Points
- Qwen3.8-Omni-Flash supports text, image, audio, and video input with a 1M context window.
- Average benchmark improvement of over 26% across 30 evaluations compared to the previous generation.
- Significant gains in Audio-Visual Agent, Coding, and long-term tasks.
- API pricing for audio input dropped by over 98%; audio/video input down by over 93%.
- Open-sourced Qwen-Live Harness and expanded Qwen-MM-Plugins for long workflows.
- Model used in its own R&D, cutting Sichuan dialect error rate by ~40.7%.
- Realtime version introduces "sound localization" capability.
With these updates, Qwen is clearly pushing the boundaries of what multimodal models can do—and making them more affordable in the process.