Grok 4.7 vs MiMo V2.6 Flash: Benchmark Comparison

Grok 4.7 has the higher score on 13 of 20 shared benchmarks; MiMo V2.6 Flash leads on 5.

The largest observed score gap is 37.00 pts on ProofBench v1.1 , where MiMo V2.6 Flash leads.

Reported ±1 standard-error ranges overlap on 12 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.7 MiMo V2.6 Flash Gap Reported uncertainty
Vals Index 54.95% ±1.07 53.23% ±1.10 1.72 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 12.50% ±2.66 11.25% ±2.54 1.25 pts Reported ±1 SE ranges overlap
Legal Research Bench 47.12% ±3.47 37.98% ±3.37 9.13 pts Reported ±1 SE ranges do not overlap
EMB 66.99% ±3.04 65.46% ±2.71 1.53 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 52.25% ±0.35 56.28% ±0.46 4.03 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 65.60% ±3.24 59.90% ±3.13 5.69 pts Reported ±1 SE ranges overlap
MedCode 49.55% ±2.17 41.06% ±2.03 8.50 pts Reported ±1 SE ranges do not overlap
MedScribe 89.38% ±1.89 85.28% ±1.99 4.10 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 26.00% ±4.41 63.00% ±4.85 37.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 69.26% ±1.33 69.26% ±0.98 0.00 pts Reported ±1 SE ranges overlap
MysteryMechanism 25.23% ±2.92 21.62% ±2.77 3.60 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 10.00% ±3.61 4.29% ±2.44 5.71 pts Reported ±1 SE ranges overlap
SAGE 40.79% ±3.35 43.53% ±3.43 2.74 pts Reported ±1 SE ranges overlap
Code Migration 44.82% ±4.21 40.93% ±4.35 3.89 pts Reported ±1 SE ranges overlap
IOI 57.72% ±1.99 47.67% ±4.79 10.05 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 28.79% ±2.31 24.24% ±1.51 4.55 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 86.17% ±2.18 78.96% ±4.04 7.22 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 69.46% ±5.67 75.36% ±5.42 5.89 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 65.63% ±1.24 67.59% ±1.22 1.96 pts Reported ±1 SE ranges overlap

Performance by category

Category Grok 4.7 average MiMo V2.6 Flash average
Legal 29.81% 24.62%
Finance 61.61% 60.55%
Healthcare 69.47% 63.17%
Math 26.00% 63.00%
Science 34.83% 31.72%
Education 40.79% 43.53%
Coding 43.60% 38.46%
Cyber 69.46% 75.36%
Social Mobility 65.63% 67.59%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.7 cost MiMo V2.6 Flash cost Grok 4.7 latency MiMo V2.6 Flash latency
Vals Index $12.12 $0.20 35m31s 58m02s
Harvey's Legal Agent Benchmark $11.13 $0.09 42m05s 17m12s
Legal Research Bench $4.92 $0.07 21m14s 15m20s
EMB $6.48 $0.11 34m57s 25m22s
Finance Agent (v2) $2.72 $0.07 15m33s 7m02s
Tax Agent Bench $1.80 $0.05 16m36s 17m19s
MedCode N/A N/A 3m33s 93.37s
MedScribe N/A N/A 2m33s 74.48s
ProofBench v1.1 $0.79 $0.12 12m47s 1h03m
BioMysteryBench $2.49 $0.05 11m37s 28m28s
MysteryMechanism $1.84 $0.06 17m33s 35m08s
Terminal-Bench Science $14.45 $0.24 1h29m 3h59m
SAGE N/A N/A 4m51s 2m16s
Code Migration $36.55 $0.49 1h09m 3h12m
IOI $12.71 $0.22 46m48s 1h49m
ProgramBench $46.49 $0.47 3h38m 3h22m
Terminal-Bench 4.0 $18.09 $0.22 50m21s 2h25m
Vibe Code Bench v1.1 $15.83 $0.56 35m47s 50m13s
CyberBench v1.1 $8.54 $0.05 15m30s 28m22s
Public Benefits Bench v1.1 $2.25 $0.03 16m57s 20m56s

Results available only for Grok 4.7

  • Vals RSI Index
  • LegalBench

Results available only for MiMo V2.6 Flash

  • SRE Bench
Model details Grok 4.7 Model details MiMo V2.6 Flash