Sep 22, 2026
Xiaomi's MiMo V2.6 Pro and MiMo V2.6 Flash evaluated across our benchmark suite
We evaluated Xiaomiโs new open-weight MiMo V2.6 Pro and MiMo V2.6 Flash through the official Xiaomi API.
-
MiMo V2.6 Flash places #16 of 65 on the Vals Index (59.58%) at $0.20 per test, with MiMo V2.6 Pro right behind at #17 (59.47%) at $0.39 per test. Both are the top-scoring open-weight models on the Index, ahead of DeepSeek V4.1 Flash (57.86%) and Kimi K3 (57.81%), and both land within a point of Grok 4.7 (60.22%) and GPT-5.6 Luna (59.88%). Flash is the cheapest model in the Indexโs top 20.
-
A large step up from the previous generation: MiMo V2.5 Pro and MiMo V2.5 scored 40.97% and 39.91% on the Vals Index.
-
Both models lead CyberBench: Flash is #1 of 9 (75.36%) and Pro #3 (72.86%), with both at 85.71% on the Patch track and Flash ahead on PoC (65.00% versus 60.00%). Proโs other strong results are on agentic tasks: #5 of 42 on Public Benefits Bench (68.95%), #9 of 103 on Vibe Code Bench (85.22%), #10 of 34 on Terminal-Bench 4.0 (24.75%), #11 of 68 on Finance Agent v2 (57.34%), #11 of 68 on Legal Research Bench (47.12%, matching Grok 4.7) and #11 of 69 on Harveyโs Legal Agent Benchmark (10.83%). Mid-board on ProofBench v1.1 (70.00%, #14 of 42), Tax Agent Bench (64.94%, #13 of 28), Code Migration (43.01%, #20 of 68) and IOI (39.33%, #26 of 33).
-
Flash beats Pro on several boards despite costing half as much: #16 of 73 on Terminal-Bench 2.1 (76.40%, matching GPT 5.5, versus Proโs 67.79% at #33), #8 of 69 on Harveyโs Legal Agent Benchmark (11.25%), #18 of 65 on EMB (65.46%, versus Proโs 62.86% at #23) and #12 of 16 on MysteryMechanism (21.62%, versus Proโs 15.32% at #15). On Terminal-Bench 2.1, Proโs failed attempts are dominated by long reasoning turns that run out the task clock, where Flash takes more, shorter turns and finishes.
-
Weak spots: both models struggle on Terminal-Bench Science (Pro 2.86%, #21 of 30; Flash 5.71%, #15), Pro fully resolves 1 of 200 tasks on ProgramBench (0.50%, #14 of 52 with a 69.82% partial pass rate), and Flash drops to #28 of 68 on Legal Research Bench (37.98%) and #20 of 28 on Tax Agent Bench (59.90%). Both models hit the 128k output cap on a handful of long agentic tasks (most often on Code Migration), which count as failures.
MiMo V2.6 Pro is priced at $0.435 per million input tokens and $0.87 per million output tokens; MiMo V2.6 Flash at $0.14 and $0.28. Both have a 1M-token context window and 128k max output tokens. We evaluated them with reasoning enabled and temperature set to 1.0. We observed no refusals across the Vals Index runs.
Congrats to the Xiaomi team on the release!