Xiaomi open-sources MiMo-V2.6-Pro and Flash, claiming 46.32 on the Artificial Analysis Intelligence Index, the top open-source score
- Xiaomi open-sourced the MiMo-V2.6 series: MiMo-V2.6-Pro and the cheaper MiMo-V2.6-Flash, both natively omnimodal, with Pro scoring 46.32 on the Artificial Analysis Intelligence Index v4.3, ahead of Kimi K3 and Qwen3.8 Max by Xiaomi's account, and API pricing unchanged from the V2.5 series.
- The RL training run was streamed live: in under six days each model completed 30 RL steps over roughly 750k trajectories, at about $0.85M for Flash and $2.62M for Pro, with relative pass rate gains of 25% and 12% on the training tasks.
- On held-out DeepSWE v1.1, scores rose about 17 points for Flash (48.8 to 65.68) and about 14 points for Pro (58.4 to 72.57), and Xiaomi says the RL gains kept improving through the run and generalized beyond the training distribution.
- MiMo-V2.6-Flash is 309B total with 15B activated parameters and MiMo-V2.6-Pro is 1.02T total with 42B activated per the Hugging Face releases; a Pro-UltraSpeed variant claims up to 20x faster output at the same quality.
- The series leads CyberGym (Flash 95.1, Pro 94.0) and Automation Bench v1.0.6 (Pro 53.1), while Claude Opus 5 still tops most other listed benchmarks, including Terminal Bench 4.0 at 59.6 against Pro's 34.9.
Hacker News opinions
Finally a lab that doesn't cheat on the charts.
The Chinese labs have gotten really good at advertising model releases. The moat is thin. What I like here: they demo diverse tasks like driving a DAW, plot benchmarks against price, and show the model used in real scientific work.
Big week. Probably the next OpenAI and Anthropic models, Grok 4.7, Mimo. These open releases are why I can't take the slow down crowd seriously. I ran older Mimo, qwen, step and gpt-oss against each other in Werewolf and Sketch.io style games, letting them trash talk while they played. Mimo was the pareto frontier for anything under $0.15/M input tokens on OpenRouter. Qwen won the shit talking though.
Cost and capability look great here, genuinely pushing lightweight open weight models forward.
So this explains why mimo 2.5 got dumber over the last two weeks. I was speculating they were about to ship a new version because the model really acted out. Now I have my proof.
Wouldn't that only be possible if your provider was Xiaomi itself?
Why do these models all love the '01 - UPPERCASE TEXT' motif in frontend designs? It's everywhere now. Cloudflare's page has '01 · QUICK TUNNELS' and no 02 anywhere.
My guess is it's scaffolding. They break sections into components, label them for themselves, and it gets reinforced by existing web patterns and by users. Probably baked in during training.
Whatever your definition of truly open is, the transparency is what stands out. The live RL dashboard was a real learning and teaching tool for me, and the tech report has the kind of behind the scenes tricks you see in DeepSeek and Google writeups. They even publish the benchmarks they did badly on.
What can you actually see in that dashboard that a casual observer can't? The metrics tab is absurdly detailed and I don't know what to make of it.
Maybe this is why US labs all sing the same slow down tune. They're worried a good enough Chinese model kills their margin, and we already have stories of US companies moving tasks to cheaper Chinese models on neoclouds.
Worth noting MiMo is led by Luo Fuli, ex-Alibaba and DeepSeek. That's probably why the tech and go-to-market look so DeepSeek shaped.
The RL dashboard also takes some air out of Anthropic's distillation attack claim, since they clearly have their own RL environments. Caveat is the RL datasets are still opaque, so nothing is really proved.
Parameters from Hugging Face: Flash is 309B total with 15B active, Pro is 1.02T total with 42B active. More like 500B in FP8, and the HF pill numbers are off.
All of this is a tease with 128GB of shared memory. Going to 256GB is a mortgage payment and it's getting tempting.
There's a Qwen 3.5 9B distill, that's an option.
They mixed up DeepSeek 4.1 Flash with something else on this page. I think that entry is actually Gemini 3.8 Flash.
Anyone know the unnamed model sitting on the pareto frontier chart between MiMo 2.5 and 2.6? Weird to acknowledge someone at the front edge and not name them.
Pretty sure that's Luna xhigh.
I really liked MiMo 2.5, cheap and it actually had vision unlike DeepSeek. Tried 2.6 Flash on a niche topic I specialise in and it did a good job. They've clearly polluted the training data with claudeslop, but past the slop there's a decent model.
How do you even recognize claudeslop?
The problem is if claudeslop infects every new model it compounds, model produces slop, slop goes into the next one. At some point we lose reliable ways to establish truth. Feels like epistemic collapse and I thought it would take longer.
Calling it a Pareto line is wrong. Pareto is the 80/20 rule. What they plotted is the frontier line, the set of models not strictly dominated by something cheaper and smarter.
Pareto front is a standard term and that's exactly what this is.
Two different concepts named after the same person. Pareto efficiency and Pareto curves are the best tradeoff along the axes, the Pareto principle is 80/20.