Alibaba's Qwen3.8-Omni-Flash cuts video input costs about 89% versus Qwen3.5-Omni-Plus
- Alibaba's Tongyi Qianwen released Qwen3.8-Omni-Flash, its first omni-modal model designed around agent capabilities, combining native audio-video understanding, reasoning and tool calling so one model handles understanding, task planning, tool calls and final output.
- Video input costs fall about 89% against Qwen3.5-Omni-Plus, and the model's agent perception mode on OmniVideoBench cut token consumption 51.8% versus static understanding.
- The model carries a 1 million token context and can locate key segments in long videos on its own, with audio-video quality described in the official announcement as close to Gemini 3.8 Flash.
- Agent performance rose an average of 19.5 points across the WildClawBench-MM and UniClawBench tests, according to the announcement.
- Alibaba open-sourced Qwen-MM-Plugins alongside the release and says Qwen-Live Harness will follow later.