Alibaba has expanded its artificial intelligence model library through the Qwen3.8-Omni-Flash multi-modal system. Within a single architecture, it can handle text, images, audio and video, and supports 1 million tokens of contextual windows. Developers pointed out that the new product has shown good results in various scenarios such as video editing, music video production, film production, aggregation of audio-visual information and real-time conversations.
Image source: qwen.ai
One of the main achievements is the reduction in service costs. Compared with the previous version of Qwen3.5-Omni-Plus, when accessed through API, the hourly audio processing price dropped by more than 98%, and the hourly audio and video processing price dropped by more than 93%. The average score of 29 benchmark tests increased by more than 25% compared with the previous generation. In tests covering audio and video agent work, writing code and performing long-term planning tasks, the model scored 71.0 points in WildClawBench-MM, 36.5 points higher than its predecessor. The UniClawBench score is 69.6. The model scored 82.7 in the LongAudioSpan test, which understands long recordings, 63.4 in OmniVideoBench, which analyzes audio and video, 28.2 in OmniCap-IF, which follows instructions when describing videos, and 89.7 in Chinese transcribing a multi-person Chinese conference.
In terms of audio and video joint processing quality, Qwen3.8-Omni-Flash is close to Google Gemini 3.8 Flash, and even surpasses its American competitors when processing only audio. The speech recognition system supports 74 languages and dialects, including Cantonese, Uyghur and Maori, and the generation system supports 29 languages. Both systems support Chinese, English, Korean, French, German, Spanish and Japanese.
One of the main features of the model is the “agent understanding” mechanism for audio and video. Instead of analyzing hours-long recordings from start to finish, the model decides “where to look and what to listen” based on the nature of the request, gradually narrowing the search in several stages. In OmniVideoBench, this mechanism helped improve results from 63.4 points to 67.8 points and reduce the number of tokens spent per request from 145,736 to 79,117, a reduction of 45.7%. Alibaba emphasized that modern solutions cannot handle long-term audio and video recordings well: the software tools of artificial intelligence agents do not yet have full built-in support for these formats, and the corresponding mechanisms are only in the early stages of development.
Alibaba also conducted an experiment using Qwen3.8-Omni-Flash to develop another AI model. Her mission is to improve the recognition of Chinese Sichuan dialect within 12 hours using the more compact Qwen2.5-Omni-3B model. “Advanced” models independently select test samples, measure initial metrics, and generate a set of training data. After four series of experiments, she created 3413 training samples, and the single character recognition error rate dropped from 25.79% to 15.30%, which is a 40.7% drop. Therefore, AI agents can not only use other models, but also improve them independently.
In addition to the model itself, the developers have also updated the accompanying software. The Qwen-MM-Plugins set for processing audio and video includes contextual search tools for media data, linking external tools, and programmatic processing of long audio and video files. Supports various environments for running AI agents, including OpenAI Codex, Anthropic Claude Code, Alibaba Qwen Code and Google Gemini CLI.
The company also released open source Qwen-Live Harness, a framework for creating interactive applications based on the Qwen3.8-Omni-Flash-Realtime API. It supports WebSocket and WebRTC protocols and provides low latency. According to Alibaba, this is the first multi-modal model with “sound localization” function: combining spatial audio data with visual information, the direction and distance of the sound source can be determined.
Qwen3.8-Omni-Flash creates clips based on musical compositions, prepares translations of short plays and commentary for feature films, compiles PDF summaries of training videos, and masters new skills through video demonstrations. During the translation process, the model can preserve individual features of the original speech. Qwen3.8-Omni-Flash is available through Alibaba’s enterprise-level artificial intelligence platform Qianwen AI.
If you find an error, select it with your mouse and press CTRL+ENTER.










