Grok 4.7 vs MiMo V2.6 Pro: Benchmark Comparison

Grok 4.7 has the higher score on 10 of 19 shared benchmarks; MiMo V2.6 Pro leads on 7.

The largest observed score gap is 44.00 pts on ProofBench v1.1 , where MiMo V2.6 Pro leads.

Reported ±1 standard-error ranges overlap on 12 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.7 MiMo V2.6 Pro Gap Reported uncertainty
Vals Index 54.95% ±1.07 55.20% ±1.18 0.25 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 12.50% ±2.66 10.83% ±2.45 1.67 pts Reported ±1 SE ranges overlap
Legal Research Bench 47.12% ±3.47 47.12% ±3.47 0.00 pts Reported ±1 SE ranges overlap
EMB 66.99% ±3.04 62.86% ±3.14 4.13 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 52.25% ±0.35 57.34% ±0.57 5.09 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 65.60% ±3.24 64.94% ±3.21 0.66 pts Reported ±1 SE ranges overlap
MedCode 49.55% ±2.17 44.97% ±2.10 4.58 pts Reported ±1 SE ranges do not overlap
MedScribe 89.38% ±1.89 88.31% ±1.94 1.07 pts Reported ±1 SE ranges overlap
ProofBench v1.1 26.00% ±4.41 70.00% ±4.61 44.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 25.23% ±2.92 15.31% ±2.42 9.91 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 10.00% ±3.61 2.86% ±2.01 7.14 pts Reported ±1 SE ranges do not overlap
SAGE 40.79% ±3.35 45.05% ±3.40 4.26 pts Reported ±1 SE ranges overlap
Code Migration 44.82% ±4.21 43.01% ±4.32 1.80 pts Reported ±1 SE ranges overlap
IOI 57.72% ±1.99 39.33% ±2.41 18.39 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 28.79% ±2.31 31.31% ±3.07 2.52 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 86.17% ±2.18 85.22% ±3.39 0.95 pts Reported ±1 SE ranges overlap
CyberBench v1.1 69.46% ±5.67 72.86% ±5.50 3.39 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 65.63% ±1.24 68.94% ±1.20 3.32 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Grok 4.7 average MiMo V2.6 Pro average
Legal 29.81% 28.97%
Finance 61.61% 61.71%
Healthcare 69.47% 66.64%
Math 26.00% 70.00%
Science 17.61% 9.09%
Education 40.79% 45.05%
Coding 43.60% 39.88%
Cyber 69.46% 72.86%
Social Mobility 65.63% 68.94%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.7 cost MiMo V2.6 Pro cost Grok 4.7 latency MiMo V2.6 Pro latency
Vals Index $12.12 $0.41 35m31s 1h07m
Harvey's Legal Agent Benchmark $11.13 $0.22 42m05s 25m49s
Legal Research Bench $4.92 $0.18 21m14s 30m16s
EMB $6.48 $0.33 34m57s 55m14s
Finance Agent (v2) $2.72 $0.20 15m33s 10m48s
Tax Agent Bench $1.80 $0.13 16m36s 25m48s
MedCode N/A N/A 3m33s 2m57s
MedScribe N/A N/A 2m33s 2m53s
ProofBench v1.1 $0.79 $0.25 12m47s 1h02m
MysteryMechanism $1.84 $0.15 17m33s 44m00s
Terminal-Bench Science $14.45 $0.62 1h29m 3h12m
SAGE N/A N/A 4m51s 3m40s
Code Migration $36.55 $0.69 1h09m 2h49m
IOI $12.71 $0.84 46m48s 2h26m
ProgramBench $46.49 $0.73 3h38m 3h00m
Terminal-Bench 4.0 $18.09 $0.50 50m21s 2h42m
Vibe Code Bench v1.1 $15.83 $1.04 35m47s 1h01m
CyberBench v1.1 $8.54 $0.09 15m30s 27m41s
Public Benefits Bench v1.1 $2.25 $0.08 16m57s 37m14s

Results available only for Grok 4.7

  • Vals RSI Index
  • LegalBench
  • BioMysteryBench

Results available only for MiMo V2.6 Pro

  • SRE Bench
Model details Grok 4.7 Model details MiMo V2.6 Pro