Xiaomi MiMo-v2.6-Pro tops open-weights Intelligence Index at 46 amid Anthropic distillation allegation; Artificial Analysis profiles detail Index v4.3.2 for MiMo and Claude Opus…
Artificial Analysis's scraped profiles of MiMo-v2.6-Pro and Claude Opus 5.5 lay out the shared ten-eval Intelligence Index v4.3.2 plus cost, token, openness, and context-window metrics but contain no numeric scores; Xiaomi's 1.02T-parameter open-weights…
Xiaomi's open-weights, natively omnimodal MiMo-v2.6-Pro (1.02 trillion total parameters, 42 billion active, MoE, MIT license) — released alongside MiMo-v2.6-Flash and a Pro-UltraSpeed mode claiming up to 20x faster generation at the same quality — ranks first among open-weights models on Artificial Analysis's Intelligence Index v4.3.2 with a score of 46, ahead of Kimi K3 and Qwen, at cited pricing of $0.435 per million input tokens and $0.87 per million output tokens. Gains are attributed to a scaled reinforcement-learning run spanning coding, agent, visual, and cyber tasks — roughly 130 hours, 75 billion tokens, and about $2.6M ($2.62M per The Decoder; Latent Space's headline rounds this to $3M) — which lifted DeepSWE from 58.4 to 72.6. Xiaomi open-sourced an RL toolkit/environment with roughly 7,000 auto-graded tasks including cybersecurity tasks drawn from OSS-Fuzz, though sources partly disagree on availability: The Decoder says the toolkit with ~7,000 tasks was open-sourced, while Latent Space says the complete 7,000-plus task datasets are not yet public (recipes and environment code slated for release). Separately, Anthropic's threat report (case GTG-16008) tracks over 400,000 exchanges routing MiMo conversations through OpenClaw and OpenCode to Claude, which Anthropic characterizes as illegal distillation; Xiaomi's response is not described in these reports. The scraped Artificial Analysis pages for both MiMo-v2.6-Pro (2026-09-22T04:02Z) and Claude Opus 5.5 (2026-09-22T16:51Z) describe the shared Intelligence Index v4.3.2 methodology, which aggregates ten evaluations — AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1 — with coverage of agentic work, coding, document reasoning, and hallucination; AA-Omniscience scores knowledge reliability and penalizes hallucinations without punishing refusals. The MiMo page adds per-task output-token counts, cost views broken into cache-hit, input, reasoning, and output token prices, and an Openness Index rating weight availability and commercial-use limits on a 0-100 scale; the Claude Opus 5.5 page adds context-window comparisons alongside per-task cost, token, and cache/output pricing. Neither scraped page itself includes numeric scores or prices — the 46 score and $0.435/$0.87 pricing come from prior ranking coverage.
- MiMo-v2.6-Pro: 1.02T total / 42B active parameters, MoE, open-weights, natively omnimodal, MIT license; launched with MiMo-v2.6-Flash and a Pro-UltraSpeed mode claiming up to 20x faster generation at equal quality.
- Artificial Analysis Intelligence Index v4.3.2: MiMo-v2.6-Pro scores 46, first among open-weights models, ahead of Kimi K3 and Qwen; cited pricing $0.435 per million input tokens and $0.87 per million output tokens.
- RL run behind the gains: roughly 130 hours, 75B tokens, about $2.6M ($2.62M per The Decoder; Latent Space headline rounds to $3M); DeepSWE improved from 58.4 to 72.6.
- RL toolkit with ~7,000 auto-graded tasks including OSS-Fuzz-derived cybersecurity tasks; sources disagree — The Decoder says the toolkit was open-sourced, Latent Space says the complete 7,000+ task datasets are not yet public, with recipes…
- Anthropic threat report GTG-16008: over 400,000 exchanges routing MiMo conversations through OpenClaw and OpenCode to Claude, characterized as illegal distillation; Xiaomi's response is not described in the reports.
- Intelligence Index v4.3.2 aggregates ten evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.
- AA-Omniscience scores knowledge reliability and penalizes hallucinations without punishing refusals.
- MiMo-v2.6-Pro AA page (scraped 2026-09-22): covers agentic work, coding, document reasoning, hallucination, per-task output tokens, cache-hit/input/reasoning/output token pricing, and a 0-100 Openness Index for weight availability and…
Coverage timelineoldest first · each row is one article
- · 5d agoMiMo-v2.6-Pro: Intelligence, Performance and Price Analysis
Hacker News · AI· 43
Artificial Analysis profiles MiMo-v2.6-Pro on intelligence benchmarks, capability scores, openness, and per-task token cost.
- · 4d agoClaude Opus 5.5 Intelligence, Performance and Price Analysis
Hacker News · AI· 55
Artificial Analysis profiles Claude Opus 5.5 on intelligence, token cost, and context window.