Alibaba releases Qwen3.8-Omni-Flash with 1M-token context, claiming Gemini 3.8 Flash-level audio and video
- Alibaba added Qwen3.8-Omni-Flash to its Qwen series, a native omnimodal model with a 1 million token context window that keeps text performance level with text-only models of the same size.
- Alibaba says the model averages more than 25% higher across 29 benchmarks than Qwen3.5-Omni-Plus, while per-hour API cost drops over 98% for voice input and over 93% for voice and video input.
- Agentic scores reach 71.0 on WildClawBench-MM (up 36.5 points from Qwen3.5-Omni-Plus) and 69.6 on UniClawBench, with 82.7 on LongAudioSpan, 63.4 on OmniVideoBench, 28.2 on OmniCap-IF and 89.7 on Chinese multi-speaker AliMeeting.
- Speech recognition covers 74 languages and speech generation 29, and Alibaba claims overall audio performance surpasses Gemini 3.8 Flash while video performance comes close to it.
- Alibaba extended Qwen-MM-Plugins and open-sourced Qwen-Live Harness built on the Qwen3.8-Omni-Flash-Realtime API, though the harness GitHub page returned a 404 at the time of writing.