Alibaba's Qwen3.8-Omni-Flash takes on Gemini 3.8 Flash with a 1M-token omnimodal window and audio input prices cut over 98%
- Alibaba launched Qwen3.8-Omni-Flash, a native omnimodal model that takes text, image, audio, and video with a 1M-token context window while holding text quality comparable to a text-only model of the same size.
- Across 29 benchmarks the model averages more than 25% above Qwen3.5-Omni-Plus, and Alibaba cut the API price per hour of audio input by over 98% and per hour of audio-visual input by over 93%.
- Agentic scores jump by 36.5 points on WildClawBench-MM and 22.3 on AgenticVBench, hitting 69.6 on UniClawBench, while AliMeeting DER and cpWER drop from 88.11 / 89.61 to 3.35 / 17.18.
- Alibaba claims audio-visual performance close to Gemini 3.8 Flash and overall audio performance above it, and shipped Qwen-MM-Plugins plus the open-sourced Qwen-Live Harness runtime for long-form audio and video work.
- The post carries a draft label and positions omni models as moving from understanding media to planning tasks, calling tools, and delivering finished work such as video editing and film commentary.
Hacker News opinions
Curious if or when we get a Qwen4 series. One thing I really like about Qwen is the huge range of model sizes, so I can mess around with tiny local llms.
I doubt Qwen3.8-Omni-X ever ships. Last open one was Qwen3-Omni-30B-A3B, and now we get Qwen3.8 27B plus a 125B most people can't run. They're clearly slowing down open weight releases.
3.8 Max is the most grounded model I've used. Talks like a normal person, doesn't wig out and start doing things on its own like Gemini does. But it's slow as hell, only available from Alibaba, and the token plan is stingy.
What would 'RL it to oblivion' actually mean in this context?
They're doing something different with the Qwen4 architecture they demoed in Flash-Next. The reasoning comes out bizarre. Between tool calls it writes stuff like 'the user's message is just system instructions, nothing to answer yet', then still fires the right tool call anyway.
If the benchmarks are really comparable, the price gap is enormous. Gemini is 1.5 in / 9.0 out per million, Qwen 3.8 is 0.15 / 0.47. That's a massive cost reduction for the same work.
'Audio-visual close to Gemini 3.8 Flash and overall audio exceeding it' is wild if true. Gemini's audio and multilingual handling was the whole selling point for a lot of people, and they're matching or beating the other benchmarks too.
Looks like the harness repo is already gone. The GitHub link just 404s.