Alibaba's Qwen3.8-Omni-Flash takes text, image, audio and video in one model, claims Gemini 3.8 Flash parity
- Alibaba's Qwen team shipped Qwen3.8-Omni-Flash on September 14, 2026: one model that accepts text, image, audio and video in the same API request, carries a 1M-token context window, and reports a 25%+ average gain over Qwen3.5-Omni-Plus across 29 evaluations.
- The model is not open weight. It ships only through Alibaba Cloud's DashScope API, with no downloadable checkpoint found on Hugging Face at publication time, unlike other Qwen and DeepSeek releases.
- API pricing for audio and audio-visual input drops 93 to 98 percent, and a separate Realtime variant streams audio and video in and synthesized speech out over WebSocket or WebRTC for live sessions.
- On OmniVideoBench, agentic mode scores 67.8 accuracy at 79,117 tokens per query versus 63.4 accuracy at 145,736 tokens for static mode, a 45.7 percent token cut with higher accuracy. Gemini 3.8 Flash still leads in absolute terms at 65.2 static and 70.1 agentic.
- Speech support covers 74 recognition languages plus 39 Chinese dialects, and generation covers 29 languages plus 7 dialects, with OpenAI-compatible Chat Completions and Responses APIs.