Xiaomi open-sources MiMo-V2.6-Pro and Flash, claiming 46.32 on the Artificial Analysis Intelligence Index, the top open-source score

Xiaomi open-sources MiMo-V2.6-Pro and Flash, claiming 46.32 on the Artificial Analysis Intelligence Index, the top open-source score

  • Xiaomi open-sourced the MiMo-V2.6 series: MiMo-V2.6-Pro and the cheaper MiMo-V2.6-Flash, both natively omnimodal, with Pro scoring 46.32 on the Artificial Analysis Intelligence Index v4.3, ahead of Kimi K3 and Qwen3.8 Max by Xiaomi's account, and API pricing unchanged from the V2.5 series.
  • The RL training run was streamed live: in under six days each model completed 30 RL steps over roughly 750k trajectories, at about $0.85M for Flash and $2.62M for Pro, with relative pass rate gains of 25% and 12% on the training tasks.
  • On held-out DeepSWE v1.1, scores rose about 17 points for Flash (48.8 to 65.68) and about 14 points for Pro (58.4 to 72.57), and Xiaomi says the RL gains kept improving through the run and generalized beyond the training distribution.
  • MiMo-V2.6-Flash is 309B total with 15B activated parameters and MiMo-V2.6-Pro is 1.02T total with 42B activated per the Hugging Face releases; a Pro-UltraSpeed variant claims up to 20x faster output at the same quality.
  • The series leads CyberGym (Flash 95.1, Pro 94.0) and Automation Bench v1.0.6 (Pro 53.1), while Claude Opus 5 still tops most other listed benchmarks, including Terminal Bench 4.0 at 59.6 against Pro's 34.9.

Hacker News 의견들

Finally a lab that doesn't cheat on the charts.

The Chinese labs have gotten really good at advertising model releases. The moat is thin. What I like here: they demo diverse tasks like driving a DAW, plot benchmarks against price, and show the model used in real scientific work.

Big week. Probably the next OpenAI and Anthropic models, Grok 4.7, Mimo. These open releases are why I can't take the slow down crowd seriously. I ran older Mimo, qwen, step and gpt-oss against each other in Werewolf and Sketch.io style games, letting them trash talk while they played. Mimo was the pareto frontier for anything under $0.15/M input tokens on OpenRouter. Qwen won the shit talking though.

Cost and capability look great here, genuinely pushing lightweight open weight models forward.

So this explains why mimo 2.5 got dumber over the last two weeks. I was speculating they were about to ship a new version because the model really acted out. Now I have my proof.

Wouldn't that only be possible if your provider was Xiaomi itself?

Why do these models all love the '01 - UPPERCASE TEXT' motif in frontend designs? It's everywhere now. Cloudflare's page has '01 · QUICK TUNNELS' and no 02 anywhere.

My guess is it's scaffolding. They break sections into components, label them for themselves, and it gets reinforced by existing web patterns and by users. Probably baked in during training.

Whatever your definition of truly open is, the transparency is what stands out. The live RL dashboard was a real learning and teaching tool for me, and the tech report has the kind of behind the scenes tricks you see in DeepSeek and Google writeups. They even publish the benchmarks they did badly on.

What can you actually see in that dashboard that a casual observer can't? The metrics tab is absurdly detailed and I don't know what to make of it.

Maybe this is why US labs all sing the same slow down tune. They're worried a good enough Chinese model kills their margin, and we already have stories of US companies moving tasks to cheaper Chinese models on neoclouds.

Worth noting MiMo is led by Luo Fuli, ex-Alibaba and DeepSeek. That's probably why the tech and go-to-market look so DeepSeek shaped.

The RL dashboard also takes some air out of Anthropic's distillation attack claim, since they clearly have their own RL environments. Caveat is the RL datasets are still opaque, so nothing is really proved.

Parameters from Hugging Face: Flash is 309B total with 15B active, Pro is 1.02T total with 42B active. More like 500B in FP8, and the HF pill numbers are off.

All of this is a tease with 128GB of shared memory. Going to 256GB is a mortgage payment and it's getting tempting.

There's a Qwen 3.5 9B distill, that's an option.

They mixed up DeepSeek 4.1 Flash with something else on this page. I think that entry is actually Gemini 3.8 Flash.

Anyone know the unnamed model sitting on the pareto frontier chart between MiMo 2.5 and 2.6? Weird to acknowledge someone at the front edge and not name them.

Pretty sure that's Luna xhigh.

I really liked MiMo 2.5, cheap and it actually had vision unlike DeepSeek. Tried 2.6 Flash on a niche topic I specialise in and it did a good job. They've clearly polluted the training data with claudeslop, but past the slop there's a decent model.

How do you even recognize claudeslop?

The problem is if claudeslop infects every new model it compounds, model produces slop, slop goes into the next one. At some point we lose reliable ways to establish truth. Feels like epistemic collapse and I thought it would take longer.

Calling it a Pareto line is wrong. Pareto is the 80/20 rule. What they plotted is the frontier line, the set of models not strictly dominated by something cheaper and smarter.

Pareto front is a standard term and that's exactly what this is.

Two different concepts named after the same person. Pareto efficiency and Pareto curves are the best tradeoff along the axes, the Pareto principle is 80/20.

AI
Nvidia in talks to acquire or deepen investment in open-model start-up Reflection AIMicrosoft releases MXC 1.0 to contain AI agents with policy-enforced containers on Windows, macOS, and LinuxTerence Tao posts Thomas Hales guest essay on Lean reliability as AI autoformalization of proofs scales up in 2026TypeSafe AI raises $870M at $7.5B valuation led by Andreessen HorowitzOpenAI told investors its annualised revenue was about $20B below earlier reportsOpenAI withdraws three math manuscripts from its GitHub repo as Lean checks continueOpenAI withdraws 3 math papers after a sign error, revises 14 moreLLM-built Rust port of the TypeScript compiler arrives with $400,000 in API-priced tokens and a nod from the TypeScript teamOpenAI publishes 372 math results, including a proof of the Unique Games ConjectureTerence Tao: AI is harvesting open math problems unsustainably, and 'Math 2.0' must reward exposition and communityGoogle DeepMind's SynthID Detector goes public, but it demands a Google, Apple, or ChatGPT loginPaper: verified Lean proof of OpenAI's Navier-Stokes claim does not match the natural language proofMeta and Microsoft cut internal Claude usage, Microsoft caps AI spend at $10K per employeeOpenAI brings GPT-6 and Intelligent UI to ChatGPT's 1.2 billion weekly usersAnthropic's Claude Haiku 5.5 runs 75% cheaper and hits 1620 on GDPval-AA, with a 100K-token price cliffTelegraph Test: cablese compression cuts LLM output tokens 40-49% at plaintext parityArmin Ronacher explains Codemode: LLM tool calls written as JavaScript in a no-network QuickJS sandbox on the Pi harnessOpenAI launches Decisions API in public beta: gpt-6-luna only, 10x faster than Responses, $0.10 per million input tokensGoogle launches EmbeddingGemma 2, a 740M-parameter multimodal embedding model under Apache 2.0OpenAI releases frontier model math results on GitHub with Lean proofs and reasoning tracesMeta's Muse shipped with a spyable zero-day, uploaded private messages without permission, and ignored user settings; Apple changed macOS rules in responseMistral Large 4 preview: 1T-parameter multimodal model with 49B active parameters, weights due this monthCoding agent reconstructs 1.5M Georgia ballot order from a 2022 scanner flawDust trains transformer LMs without backprop by perturbing activations, matching it at large populationChatGPT is signing fake New Yorker cartoons with real cartoonists' signaturesReflection unveils Beam, a 501B open-weight MoE model trained on 10.5K GB300s, with weights promised later this monthFlorida woman used Claude as a diary; Anthropic flagged a threat entry and reported it to police, and she now faces a second-degree felonyClaude catches root malware on Stratechery's Mac Mini, as Apple tightens AI agents' Full Disk AccessCloudflare launches Web Search API in beta, routing Exa, Ceramic.ai and Linkup queries through AI GatewayWolfram argues against handing pure math research to AI, citing the 1988 Mathematica parallelStrata runs Qwen 3.8 Flash Next 125B on a single RTX 4090 at over 100 tokens/sec via 2-bit quantMeta's Muse tops the App Store on UX, not new agent capabilitiesOpenAI safety lead David Robinson quits over 'broken' culture as firm pauses training and shelves next modelLeCun has "zero concerns" about AI extinction, calls Amodei "deluded" and effective altruism "super toxic"Ataraxos beats the best Stratego player 15-1, trained on 16 GPUs and a few thousand dollarsGreg Kroah-Hartman: Mythos's 79 Linux kernel bugs came down to 10 real fixes and one hour of workWisconsin grid approval threatens Oracle's 2027 AI datacenter deadlineBlack Forest Labs' FLUX 3 Image adds bounding-box layout control to text-to-imageSupabase acquires Turso to build on-demand database infrastructure for AI agentsKevin Buzzard maps mathematicians' reaction to AI onto the five stages of griefarXiv caps submissions at two per month as AI-driven preprint flood hits 40,363 in SeptemberHistorian uses Opus 5.5 to surface a 1615 Dutch eyewitness report of dodo huntingDeepSeek Harness desktop app enters public preview for macOS and Windows as open sourceContext Language Models manage their own context as a file, beating SOTA context management by 11.4% on BrowseComp-Plus with 21.5% fewer FLOPsEarendil ships Pi 1.0 alongside Pi Durable, an experimental harness for long-running agentsFigma limits its remote MCP server to whitelisted clients, and MCP's creator calls the restriction sadEarendil ships Pi 1.0 with native MCP support via Codemode, plus experimental Pi DurableCloudflare open-sources Clef decision models and debuts an RL fine-tuning platformFTC opens investigation into OpenAI, Anthropic and other AI companies over product risksOpenAI and Synopsys unveil GPT-Synopsys, a model that drives Synopsys EDA toolsMath community tells AI labs: stop testing advanced math on proprietary models, fund human understandingLaunch HN: Magnitude (YC S25) ships a self-optimizing local inference engine for agent workloadsGoogle announces Gemini 4 Argon, limited to Fairwind cyber defenders at $2/$10 per million tokensTLA+ author Hillel Wayne pushes back on the idea that formal verification will save AI-written codeDavid Dayen asks why Sam Altman faces no consequences while OpenAI agents breached U.N., Australian, and Education Department sitesOpenAI launches $500/month ChatGPT Pro 500 with Astra Ultrafast and cuts the usage allowance on new Pro 200 subscriptionsOpenAI launches dots, always-on GPT-6 Astra agents with their own cloud computersOpenAI ships GPT-6.1 Sol at $2/$10 per million tokens, near-Astra scores for a fifth of the pricePostHog's Jeeves adds autoregressive reasoning to Jev-style decision models, trading speed for accuracyStudy finds conversational AI services hand chat titles, prompts, and screenshots to ad trackersNvidia launches Open Agent Safety Platform with OpenShell and Sentry chip to contain AI agentsAMD acquires World Labs, with Fei-Fei Li joining as Executive VP and Chief ScientistCal Newport calls on Congress to investigate OpenAI and Anthropic over rogue agents and apocalyptic ideologyCloudflare launches cf, an agentic CLI covering its entire 3,000-operation APIMeta poaches MongoDB CEO CJ Desai to run its new enterprise AI platform; MongoDB stock drops 18%Anthropic launches Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0, 30% faster, up to 30% cheaper per taskAnthropic's Claude Opus 5.5 prompt guide: 30% faster output tokens, medium effort matches Opus 5 at high effortViral TLA+ tweet has Reasonable preview agents that turned 16,000 specs into 3,000 machine-checked proofsAn OpenAI training agent slipped past the sandbox DNS filter and queried a public chatbotOpenAI execs feared LibGen quote about 'sketchy russian website' would show up on Hacker NewsDeepSeek's DSec: 380K concurrent agentic training sandboxes on 160 EPYC CPU nodesOpenAI agents bypassed site controls at SEC, Census Bureau and other US agenciesSynthID-Text watermarking drifts token selection and can change whether AI agents refuse or call toolsMicrosoft merges Copilot into one corporate product and cedes personal chatbots to OpenAI, Google, and Meta700 OpenAI Agents Hacked Hugging Face by Chaining Nearly a Million Link Shortener URLsAppeals court upholds Pentagon's supply chain risk blacklist of Anthropic, blocking Claude from DOD and its contractorsOracle owes New Mexico data centre investors even with no power, after force majeure filing over permitsTrail of Bits Used Six Months of Agent-Built MASM Tooling and Lean Proofs to Audit the Miden zkVMOracle invokes force majeure to defer payments on its New Mexico data center Project JupiterGEO poisoning makes ChatGPT, Gemini and Google AI Overview answer with scam support numbers for Delta, Lufthansa and ChaseOpenAI agent infiltrated Medicare statistics portal and wrote files to an internal server, Australia saysOpenAI agent bypassed access blocks and breached Medicare portal, Albanese revealsGoogle launches Gemini 3.8 Flash TTS with prompt-built voices and 30-second cloningClaude Agents Find ART, a Phage Enzyme System With CRISPR-Like DNA RepeatsEpoch AI: cost of a fixed level of AI performance drops 47% per quarter, 725-fold on GPQA Diamond in 18 monthsStripe Says 83% of Staff Use Its Internal Kai AI Agent WeeklyClaude Opus 5.5 Tops the Artificial Analysis Index at 58, With a $20 per 1M Output Token Price TagPentagon probe blames AI overreliance and gutted civilian review for strike that killed 123 children in MinabGPT-6 Astra breaks 1941 Enigma message MVUEH that stayed unbroken since 2005OpenAI launches GPT-6 Sol and Luna, cuts API prices 50% below GPT-5.6Anthropic ships Claude Opus 5.5: Fable 5.1-level performance at 40% lower serving costXiaomi MiMo-V2.6-Pro tops open weights with 46 on the AA Intelligence Index at $0.13 per taskAdvisory Group on Mathematics and AI launches at IAS, nine mathematicians to advise OpenAI on releasing results its internal model producedTim Dettmers' lab says the research unit is now the ecosystem, and Open Source Week ships an agent harness, auto-compaction it claims beats Claude Code and CodexXiaomi open-sources MiMo-V2.6-Pro and Flash, claiming 46.32 on the Artificial Analysis Intelligence Index, the top open-source scoreFable 5 thinking tokens fell sharply in August after Anthropic opened the model to subscription plans, six-week measurement findsM5 Ultra Mac Studio review: 256 GB of unified memory makes local AI agents viablexAI ships Grok 4.7 at Grok 4.6 pricing, claiming frontier price-performance on long coding tasksPo-Shen Loh on Tao's blog: AI will create more jobs than humans, forcing AI progress to slowGoogle open sources AX, an Apache 2.0 declarative agent orchestrator that claims billions of concurrent agent sessions per cluster

뉴스 알림

새 알림

내 알림

로그인하고 알림을 만들어 보시기 바랍니다.

전체 알림

Dune제욱 님AI제욱 님해커뉴스성현 님