English

Model ReleasesAlibabaQwen3.8-Omni-Flash

Alibaba Releases Qwen3.8-Omni-Flash: A High-Efficiency Multimodal Model with Enhanced Audio-Visual Capabilities

Alibaba has added Qwen3.8-Omni-Flash, a next-generation native omnimodal model, to its Qwen series. The model is designed to strengthen agent capabilities in real-world productivity scenarios, moving beyond simple multimodal understanding to planning tasks, using tools, and completing creative workflows such as video editing, music production, and real-time conversations.

Qwen3.8-Omni-Flash supports a context window of up to 1 million tokens while maintaining text performance comparable to text-only models of a similar size. In benchmark tests across 29 evaluations, the model showed an average score improvement of over 25% compared to its predecessor, Qwen3.5-Omni-Plus. Notably, the model offers significant cost advantages, with the hourly API price for audio input reduced by more than 98% and audio-visual input reduced by over 93%.

The model demonstrates high performance in specialized benchmarks, including WildClawBench-MM and UniClawBench, specifically for audio-visual agents and long-horizon tasks. Its core capabilities, such as long-form audio understanding, audio-visual reasoning, and multi-speaker recognition, have also seen substantial improvements. Alibaba stated that these advancements signify the evolution of audio and video from mere perceptual inputs into core media that allow agents to understand environments, reason, and execute complex tasks.

Sources

  1. 「Qwen3.8-Omni-Flash」リリース、Gemini 3.8 Flashに匹敵する音声・映像処理性能を達成 (GIGAZINE, 2026-09-18)
  2. Qwen 3.8 Omni Flash (Hacker News Frontpage, 2026-09-17)
1 more sourcesHide sources
  1. Alibaba Cloud