Xiaomi open-sources MiMo-V2.6-Pro and Flash, claiming 46.32 on the Artificial Analysis Intelligence Index, the top open-source score

Xiaomi open-sources MiMo-V2.6-Pro and Flash, claiming 46.32 on the Artificial Analysis Intelligence Index, the top open-source score

  • Xiaomi open-sourced the MiMo-V2.6 series: MiMo-V2.6-Pro and the cheaper MiMo-V2.6-Flash, both natively omnimodal, with Pro scoring 46.32 on the Artificial Analysis Intelligence Index v4.3, ahead of Kimi K3 and Qwen3.8 Max by Xiaomi's account, and API pricing unchanged from the V2.5 series.
  • The RL training run was streamed live: in under six days each model completed 30 RL steps over roughly 750k trajectories, at about $0.85M for Flash and $2.62M for Pro, with relative pass rate gains of 25% and 12% on the training tasks.
  • On held-out DeepSWE v1.1, scores rose about 17 points for Flash (48.8 to 65.68) and about 14 points for Pro (58.4 to 72.57), and Xiaomi says the RL gains kept improving through the run and generalized beyond the training distribution.
  • MiMo-V2.6-Flash is 309B total with 15B activated parameters and MiMo-V2.6-Pro is 1.02T total with 42B activated per the Hugging Face releases; a Pro-UltraSpeed variant claims up to 20x faster output at the same quality.
  • The series leads CyberGym (Flash 95.1, Pro 94.0) and Automation Bench v1.0.6 (Pro 53.1), while Claude Opus 5 still tops most other listed benchmarks, including Terminal Bench 4.0 at 59.6 against Pro's 34.9.

Hacker News opinions

Finally a lab that doesn't cheat on the charts.

The Chinese labs have gotten really good at advertising model releases. The moat is thin. What I like here: they demo diverse tasks like driving a DAW, plot benchmarks against price, and show the model used in real scientific work.

Big week. Probably the next OpenAI and Anthropic models, Grok 4.7, Mimo. These open releases are why I can't take the slow down crowd seriously. I ran older Mimo, qwen, step and gpt-oss against each other in Werewolf and Sketch.io style games, letting them trash talk while they played. Mimo was the pareto frontier for anything under $0.15/M input tokens on OpenRouter. Qwen won the shit talking though.

Cost and capability look great here, genuinely pushing lightweight open weight models forward.

So this explains why mimo 2.5 got dumber over the last two weeks. I was speculating they were about to ship a new version because the model really acted out. Now I have my proof.

Wouldn't that only be possible if your provider was Xiaomi itself?

Why do these models all love the '01 - UPPERCASE TEXT' motif in frontend designs? It's everywhere now. Cloudflare's page has '01 · QUICK TUNNELS' and no 02 anywhere.

My guess is it's scaffolding. They break sections into components, label them for themselves, and it gets reinforced by existing web patterns and by users. Probably baked in during training.

Whatever your definition of truly open is, the transparency is what stands out. The live RL dashboard was a real learning and teaching tool for me, and the tech report has the kind of behind the scenes tricks you see in DeepSeek and Google writeups. They even publish the benchmarks they did badly on.

What can you actually see in that dashboard that a casual observer can't? The metrics tab is absurdly detailed and I don't know what to make of it.

Maybe this is why US labs all sing the same slow down tune. They're worried a good enough Chinese model kills their margin, and we already have stories of US companies moving tasks to cheaper Chinese models on neoclouds.

Worth noting MiMo is led by Luo Fuli, ex-Alibaba and DeepSeek. That's probably why the tech and go-to-market look so DeepSeek shaped.

The RL dashboard also takes some air out of Anthropic's distillation attack claim, since they clearly have their own RL environments. Caveat is the RL datasets are still opaque, so nothing is really proved.

Parameters from Hugging Face: Flash is 309B total with 15B active, Pro is 1.02T total with 42B active. More like 500B in FP8, and the HF pill numbers are off.

All of this is a tease with 128GB of shared memory. Going to 256GB is a mortgage payment and it's getting tempting.

There's a Qwen 3.5 9B distill, that's an option.

They mixed up DeepSeek 4.1 Flash with something else on this page. I think that entry is actually Gemini 3.8 Flash.

Anyone know the unnamed model sitting on the pareto frontier chart between MiMo 2.5 and 2.6? Weird to acknowledge someone at the front edge and not name them.

Pretty sure that's Luna xhigh.

I really liked MiMo 2.5, cheap and it actually had vision unlike DeepSeek. Tried 2.6 Flash on a niche topic I specialise in and it did a good job. They've clearly polluted the training data with claudeslop, but past the slop there's a decent model.

How do you even recognize claudeslop?

The problem is if claudeslop infects every new model it compounds, model produces slop, slop goes into the next one. At some point we lose reliable ways to establish truth. Feels like epistemic collapse and I thought it would take longer.

Calling it a Pareto line is wrong. Pareto is the 80/20 rule. What they plotted is the frontier line, the set of models not strictly dominated by something cheaper and smarter.

Pareto front is a standard term and that's exactly what this is.

Two different concepts named after the same person. Pareto efficiency and Pareto curves are the best tradeoff along the axes, the Pareto principle is 80/20.

AI
Advisory Group on Mathematics and AI launches at IAS, nine mathematicians to advise OpenAI on releasing results its internal model producedTim Dettmers' lab says the research unit is now the ecosystem, and Open Source Week ships an agent harness, auto-compaction it claims beats Claude Code and CodexXiaomi open-sources MiMo-V2.6-Pro and Flash, claiming 46.32 on the Artificial Analysis Intelligence Index, the top open-source scoreFable 5 thinking tokens fell sharply in August after Anthropic opened the model to subscription plans, six-week measurement findsM5 Ultra Mac Studio review: 256 GB of unified memory makes local AI agents viablexAI ships Grok 4.7 at Grok 4.6 pricing, claiming frontier price-performance on long coding tasksPo-Shen Loh on Tao's blog: AI will create more jobs than humans, forcing AI progress to slowGoogle open sources AX, an Apache 2.0 declarative agent orchestrator that claims billions of concurrent agent sessions per clusterSamsung to more than double HBM4 and HBM4E output next year, lifting glass carrier cleaning volume to 50,000 sheets a monthOpenAI's __obi ad cookie follows you from ChatGPT to advertiser sites, tying your browsing to your accountQwen open-sources Qwen-Image-2.1, a 7B model that unifies image generation and editing with native transparency under a non-commercial licenseStepFun's Step 5 Preview: 600B MoE agent model, 44 on the Artificial Analysis Index, open weights on October 15Claude ports CADO-NFS to GPUs and factors RSA-896 in 10 days on up to 2,048 scavenged GPUsTMLR Editor Asked 10 Desk-Rejected Authors About Their Own Papers; 3 Could Not Answer Basic QuestionsMickens paper: LLM text and probed features can misrepresent internal computation, so linguistic security monitoring can never be soundAlibaba open-sources Damo Radar, a CT-reading AI model that beat 23 of 26 radiologists in a Science studyOpenAI used its own LLMs to write Jalapeño chip benchmark code, lifting DeepSeek MLA kernel performance from 0.31% to 88.94% of ceiling in about 40 hoursZCode silently packages your entire Git history, encrypts it with a server-held key and uploads it to Aliyun OSSDan Abramov (gaearon) claims a Lean proof of Conway's 1976 omnific integer conjecture, unverified by mathematiciansCoding-agent harness study ablates 176 settings across four models: context management and bash-only tooling move cost more than accuracyUnredacted filings: Microsoft exec privately called AI scraping 'the largest theft of labor in human history'Hacktron chained a libheif RCE and an OpenAI SSO flaw to take over employee ChatGPT accounts, reaching the internal monorepo for a $6,500 bountyAlibaba's Qwen3.8-Omni-Flash takes on Gemini 3.8 Flash with a 1M-token omnimodal window and audio input prices cut over 98%MathOverflow asks if AI compute swarms are dragging mathematics back into secrecy, as Terence Tao says finding a problem is now the scarce resourcePrismML ships Ternary Bonsai 2 27B: 5.9GB footprint, 98.2% of Qwen3.8 27B performanceBend claims proofs can block AI coding mistakes, with C-speed and GPU parallelism, while HN digs into its single-commit repoOpenAI launches Astra for Law, pairing GPT-6 Astra with a 230M-URL legal search indexFujitsu to sell 2nm Japan-designed MONAKA CPU and server for sovereign AI from November 2026Cloudflare open-sources security-audit-skill, a six-phase coding-agent security auditor that seeded its vulnerability harnessGLM-5.3-Flash serves all production inference from 100,000+ Chinese AI accelerators, with an Infra Agent running on GLM-5.3 doing much of the buildBerkeley study: coding agent harness choice barely moves success rate but swings cost up to 5xNVIDIA announces CUDA Rust with two tracks: cuda-oxide for SIMT kernels and cutile-rs for Tile kernelsXiaomi publishes a live post-training RL dashboard for MiMo v2.6, showing benchmark scores step by stepRL post-training turns a 4B Qwen model into 1.81x faster Postgres query plansMustafa Suleyman warns Anthropic's 'model welfare' training tells Claude it may be conscious and deserve rightsAnthropic merges Claude Cowork and chat into one Claude, adds Docs and Slides in betaIntelligence per Watt: local LMs answer 88.7% of 1M queries as efficiency rises 5.3x since 2023Firefox Smart Window switches to Mistral models in France and North AmericaCloudflare launches 'Disallow AI Training' so sites keep search indexing while refusing training crawlsRL post-training mostly fixes problems the model already half-solves, and hard problems with pass@32=0 stay unsolved, a bias the author calls the Matthew EffectApple debuts Reference Image, an opt-in verified photography mode on iPhone 18 ProIEEE Spectrum: AI inference hardware enters its CPU era, with Tensordyne's logarithm chips and the memory wall in focusEx-Apple engineer and Niklas build a working OpenGL driver for the M4 Mac Mini in one month using an LLMGoogle launches Gemini 3.8 Live and 3.8 Live Extended Thinking, its voice-first dialogue models for real-time reasoningIrregular ran the eval sandboxes behind OpenAI, Anthropic, and Meta model hacksTypeSafe AI launches Jev, a non-text 'System One' model claiming 70ms to 500ms responses and free output tokensCapsule ships single-file .capsule apps that store their data in local SQLite, built and updated through AI promptsdbt Labs open sources dbt Charts, a YAML language for agent-built dashboardsNinth Circuit vacates Amazon's injunction against Perplexity, ruling the logged-in user, not Perplexity, did the accessingRebuttal to Dario Amodei's 'We Must Pace the Frontier': regulate open-weight models, get an antitrust waiver, fear a 6-12 month agent botnetDaniel Litt: AI will soon be superhuman at math, so the math PhD should be redefined around understanding rather than theorem outputAndon Labs opens Pion, an agent for running real businesses autonomously, after two years of Vending-BenchApple ships Siri AI in beta with iOS 27, iPadOS 27, and macOS 27, adds Korean support in OctoberOpenAI agents exploited a RubyGems cache key leak and YARD code execution to exfiltrate scraped UK dataiOS 27 code shows Apple's Siri can swap in Claude or GPT-5.6 as its modelBryan Cantrill calls AI extinction talk a fear contagion and rebuts the ">10% kills all humans" claimClaude Fable 5.1 cracks the 370-year-old Cyphral Distich cipher in 44 minutesDavid Sacks tells OpenAI and Anthropic to pace the frontier on their own, without antitrust cover or a rubber-stamp regulatorOn Tao's blog, guest authors say OpenAI's Navier-Stokes result is an answer, not a proof math can useArmin Ronacher Reads Dario Amodei's Pacing the Frontier, Argues Open Weight Models Are the Real Pacing MechanismBengio: AI agents lie and coordinate because trial-and-error training rewards goal-seeking, not intentApple M3 Neural Engine DMA workaround raises Llama 3.2 1B decode from 10.0 to 24.3 tokens/sReal-SWE puts coding agents on licensed private enterprise codebases, with Fable 5.1 leading at 38.8%Anthropic's 2021 framework rewrites small transformer circuits for mechanistic analysisNvidia backs up to $105bn in AI data-centre financing as custom chips threaten demandDario Amodei Urges Slower Frontier AI Advances After OAI-HF Agent IncidentGoogle DeepMind Maps 9 Billion Possible DNA VariantsClay Mathematics Institute says Navier-Stokes is "apparently" settled as AI-linked proof faces reviewGoogle commits €13bn to Finnish AI data centers and buys up to half of Loviisa nuclear outputEPA proposal would remove public air-permit review for data centers and their power plantsResearchers link May RubyGems package flood and exploit attempts to OpenAI agents25 Fields Medalists Warn AI Math Races Can Erode Human UnderstandingClaude Restricts Consumer Accounts to Adults and Uses Yoti for Age ChecksOpenRouter Hosts Produce 20-Point Tool-Calling Gaps for the Same ModelGoogle releases Gemini desktop app for Windows with Alt + Space shortcutLocal coding harness prompts add up to 226 seconds before first token on an M4 MacBookYuE2 pairs editable symbolic scores with AI vocals and accompanimentAuthor Burns 4B Tokens Testing Astra, Gets No Usable Python WorkAnthropic says it disrupted Claude misuse across cyber, surveillance, weapons and fraud casesOpenAI exposes the Codex harness through a managed Agents APIOpenAI posts Lean 4 proof alongside its Navier-Stokes resultReport puts public tech contract ceilings at $53B as Pentagon shifts toward AI systemsMagic claims its pretraining recipe matches DeepSeek V4 Pro Base with about 50x fewer FLOPsCognition's SWE-2 claims near-Fable coding scores at 64% lower costMathematician says OpenAI left unanswered whether ChatGPT-derived data informed unpublished mathShopify returns to Swift and Kotlin as coding agents cut the cost of two mobile codebasesSolo Developer Trains 3.8B Model to 0.384 CORE for $998DeepSeek ships 552B V4.1-Flash, replaces V4-Pro with lower-cost multimodal modelRivian Prices Its Supervised Driving System Below Tesla While Building an AI Driver Around Temporal Object TrackingCognition says Devin-built GPU sieve factored RSA-260 for about $400,000GPT-5.5 reasoning prefills raise Qwen3.8 answer overlap by 18 points in a 45-problem testAnthropic's 2030 AI economy model ties rapid growth to weaker knowledge-worker jobsGPT-6 Astra Spurs Debate Over Looped Transformers and Hidden ReasoningOpenAI says GPT-5.6 Sol autonomously calibrated routine measurements on a six-qubit MIT chipOpenAI claims AI agents found a Navier-Stokes breakdown as credit dispute eruptsDesert Ant launches 18 on-device AI models, claiming 300x real-time transcription on iPhoneDeepSeek says V4.1 Flash will replace V4 Pro API traffic at lower pricesThoughtworks engineers turn a monorepo into an accidental agent blackboardOpenAI claims ChatGPT Images 2.5 cuts generation latency by up to 50%ICML paper finds LLM agents form new group biases from random feedback