Xiaomi MiMo-V2.6-Pro tops open weights with 46 on the AA Intelligence Index at $0.13 per task
- Xiaomi's MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index v4.3.2, first among the 114 models in its open weights class and far above the class median of 18.
- Pricing is $0.435 per 1M input tokens and $0.87 per 1M output tokens, with a 99% cache discount, giving $0.13 per Intelligence Index task; the full index evaluation cost $206.66.
- Technical specs: 1.0T total parameters with 42B active, 1M token context window, MIT license, weights on Hugging Face, and text, image, speech, and video input with text output.
- Throughput is 124.5 output tokens per second (12th of 114) and the model generated 140M tokens on the index, which the page calls somewhat verbose while the median it cites is also 140M.
- A Hacker News commenter reports Xiaomi's blog puts the model's RL training run at $3.47M, against $100M+ estimates for the big labs, and the MiMo-V2.6-Pro page itself says only the reasoning version is shown.
Hacker News opinions
The page says it generated 140M tokens which is 'somewhat verbose in comparison to the median of 140M'. Nice to see a human typo, at least a bot didn't just stop this together.
46 for MiMo-V2.6-Pro while DeepSeek-V4.1 gets 39 feels suspicious. Xiaomi's own appendix shows DeepSeek sometimes surpasses MiMo. A week ago Opus 5 sat a point ahead of Fable 5 despite Fable being much smarter, and that got corrected.
The main AA benchmark keeps changing, they had to radically rework it when Astra showed zero improvement over GPT-5.6 Sol. Opus 5 is still a point ahead of Fable 5.0 if you add it back manually, only Fable 5.1 shows ahead. The hard part is finding benchmarks that match your own use.
DS 4.1 Flash is good but clearly not as smart as the non-flash models, it lacks the training data. Without a solid plan it goes off the rails pretty regularly.
Ran it on my own benchmark suite and I'd disagree that it's fast. Faster than DeepSeek, still much slower than the leading models. Pricing is where it shines. KillSwitch-Bench 1.0: Opus 5 66.9, GPT-6 Astra 57.9, Fable 5.1 46.7, MiMo-V2.6-Pro 38.8, Muse Spark 1.3 36.5.
Speed swings a lot with demand. Last night I was seeing 80+ tok/s from it.
That looks like you aren't using the ultra speed endpoint.
OpenAI cut usage limits hard and the intelligence seems to be declining, so I'm finally trying these Chinese models seriously. I don't mind if it takes longer, I just need it to work the same way day to day.
I pay $200/mo for Codex and a weekly limit used to cover a week of work, now I get one or two days out of it. Staying on Sol orchestrating Luna Xhigh but it's still tight, and whenever a new model is about to ship the current one feels dumbed down.
Serious question, does anyone actually have evidence that intelligence is declining? People have asserted it since 2023 and every tracker I look at is just a flat line.
Sol 5.6 xhigh was a reliable coding workhorse on the $200 sub, but this week every model hits rate limits constantly without me being near the weekly cap. Early MiMo 2.6 Pro results are encouraging for anything non-UI, so I'm switching spend for now.
Xiaomi says the MiMo v2.6 training run cost $3.47M, versus $100M+ estimates for the big five. For a model that matches Muse Spark 1.3 in benchmarks with cache rates at $0.0036 per million, that's incredibly cheap.
I got the impression that $3.47M covers RL training only, their blog says as much. A barely-trained model is not scoring 48 on DeepSWE v1.1.
Why does the top of the page label it #1/114 for intelligence when the graphs further down clearly show it isn't?
Hover the #1/114 badge, it's filtered to open weights only. That's probably why.
Mimo 2.5 was really good but looped too much for my taste, and Pro 2.5 locally wasn't much better. It's a slept-on model, most people used it because it was free, but worth another shot for one-shots if they fixed the looping.