Grok 4.6 vs MiMo V2.6 Pro: Benchmark Comparison

Grok 4.6 has the higher score on 8 of 19 shared benchmarks; MiMo V2.6 Pro leads on 11.

The largest observed score gap is 19.00 pts on ProofBench v1.1 , where MiMo V2.6 Pro leads.

Reported ±1 standard-error ranges overlap on 10 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.6 MiMo V2.6 Pro Gap Reported uncertainty
Vals Index 52.09% ±1.15 55.20% ±1.18 3.11 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 15.83% ±2.65 10.83% ±2.45 5.00 pts Reported ±1 SE ranges overlap
Legal Research Bench 48.08% ±3.47 47.12% ±3.47 0.96 pts Reported ±1 SE ranges overlap
EMB 62.73% ±3.08 62.86% ±3.14 0.13 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.68% ±0.67 57.34% ±0.57 3.66 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 70.79% ±3.00 64.94% ±3.21 5.85 pts Reported ±1 SE ranges overlap
MedCode 44.71% ±2.26 44.97% ±2.10 0.26 pts Reported ±1 SE ranges overlap
MedScribe 86.53% ±1.96 88.31% ±1.94 1.77 pts Reported ±1 SE ranges overlap
ProofBench v1.1 51.00% ±5.02 70.00% ±4.61 19.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 30.63% ±3.10 15.31% ±2.42 15.32 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 11.43% ±3.83 2.86% ±2.01 8.57 pts Reported ±1 SE ranges do not overlap
SAGE 28.90% ±3.08 45.05% ±3.40 16.15 pts Reported ±1 SE ranges do not overlap
Code Migration 44.57% ±4.48 43.01% ±4.32 1.56 pts Reported ±1 SE ranges overlap
IOI 47.61% ±2.09 39.33% ±2.41 8.28 pts Reported ±1 SE ranges do not overlap
ProgramBench 1.00% ±0.70 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 17.17% ±1.34 31.31% ±3.07 14.14 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 76.24% ±3.82 85.22% ±3.39 8.98 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 67.68% ±5.87 72.86% ±5.50 5.18 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 66.85% ±1.23 68.94% ±1.20 2.10 pts Reported ±1 SE ranges overlap

Performance by category

Category Grok 4.6 average MiMo V2.6 Pro average
Legal 31.95% 28.97%
Finance 62.40% 61.71%
Healthcare 65.62% 66.64%
Math 51.00% 70.00%
Science 21.03% 9.09%
Education 28.90% 45.05%
Coding 37.32% 39.88%
Cyber 67.68% 72.86%
Social Mobility 66.85% 68.94%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.6 cost MiMo V2.6 Pro cost Grok 4.6 latency MiMo V2.6 Pro latency
Vals Index $4.49 $0.41 37m01s 1h07m
Harvey's Legal Agent Benchmark $4.01 $0.22 45m05s 25m49s
Legal Research Bench $1.53 $0.18 25m04s 30m16s
EMB $3.06 $0.33 52m59s 55m14s
Finance Agent (v2) $1.66 $0.20 16m41s 10m48s
Tax Agent Bench $0.98 $0.13 10m58s 25m48s
MedCode N/A N/A 2m07s 2m57s
MedScribe N/A N/A 70.94s 2m53s
ProofBench v1.1 $0.76 $0.25 14m22s 1h02m
MysteryMechanism $1.31 $0.15 27m55s 44m00s
Terminal-Bench Science $4.82 $0.62 1h18m 3h12m
SAGE N/A N/A 3m22s 3m40s
Code Migration $14.58 $0.69 1h30m 2h49m
IOI $8.18 $0.84 3h57m 2h26m
ProgramBench $16.76 $0.73 4h27m 3h00m
Terminal-Bench 4.0 $5.07 $0.50 29m27s 2h42m
Vibe Code Bench v1.1 $4.88 $1.04 25m30s 1h01m
CyberBench v1.1 $4.07 $0.09 36m52s 27m41s
Public Benefits Bench v1.1 $0.92 $0.08 25m53s 37m14s

Results available only for Grok 4.6

  • Vals RSI Index
  • LegalBench
  • MortgageTax
  • TaxEval v2
  • BioMysteryBench
  • GPQA Diamond
  • MMLU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
  • Time Horizon Index: KSP

Results available only for MiMo V2.6 Pro

  • SRE Bench
Model details Grok 4.6 Model details MiMo V2.6 Pro