Grok 4.6 vs MiMo V2.6 Flash: Benchmark Comparison

Grok 4.6 has the higher score on 10 of 20 shared benchmarks; MiMo V2.6 Flash leads on 10.

The largest observed score gap is 14.63 pts on SAGE , where MiMo V2.6 Flash leads.

Reported ±1 standard-error ranges overlap on 11 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.6 MiMo V2.6 Flash Gap Reported uncertainty
Vals Index 52.09% ±1.15 53.23% ±1.10 1.14 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 15.83% ±2.65 11.25% ±2.54 4.58 pts Reported ±1 SE ranges overlap
Legal Research Bench 48.08% ±3.47 37.98% ±3.37 10.10 pts Reported ±1 SE ranges do not overlap
EMB 62.73% ±3.08 65.46% ±2.71 2.73 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.68% ±0.67 56.28% ±0.46 2.59 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 70.79% ±3.00 59.90% ±3.13 10.88 pts Reported ±1 SE ranges do not overlap
MedCode 44.71% ±2.26 41.06% ±2.03 3.66 pts Reported ±1 SE ranges overlap
MedScribe 86.53% ±1.96 85.28% ±1.99 1.26 pts Reported ±1 SE ranges overlap
ProofBench v1.1 51.00% ±5.02 63.00% ±4.85 12.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 72.22% ±0.00 69.26% ±0.98 2.96 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 30.63% ±3.10 21.62% ±2.77 9.01 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 11.43% ±3.83 4.29% ±2.44 7.14 pts Reported ±1 SE ranges do not overlap
SAGE 28.90% ±3.08 43.53% ±3.43 14.63 pts Reported ±1 SE ranges do not overlap
Code Migration 44.57% ±4.48 40.93% ±4.35 3.65 pts Reported ±1 SE ranges overlap
IOI 47.61% ±2.09 47.67% ±4.79 0.06 pts Reported ±1 SE ranges overlap
ProgramBench 1.00% ±0.70 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 17.17% ±1.34 24.24% ±1.51 7.07 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 76.24% ±3.82 78.96% ±4.04 2.72 pts Reported ±1 SE ranges overlap
CyberBench v1.1 67.68% ±5.87 75.36% ±5.42 7.68 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 66.85% ±1.23 67.59% ±1.22 0.74 pts Reported ±1 SE ranges overlap

Performance by category

Category Grok 4.6 average MiMo V2.6 Flash average
Legal 31.95% 24.62%
Finance 62.40% 60.55%
Healthcare 65.62% 63.17%
Math 51.00% 63.00%
Science 38.09% 31.72%
Education 28.90% 43.53%
Coding 37.32% 38.46%
Cyber 67.68% 75.36%
Social Mobility 66.85% 67.59%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.6 cost MiMo V2.6 Flash cost Grok 4.6 latency MiMo V2.6 Flash latency
Vals Index $4.49 $0.20 37m01s 58m02s
Harvey's Legal Agent Benchmark $4.01 $0.09 45m05s 17m12s
Legal Research Bench $1.53 $0.07 25m04s 15m20s
EMB $3.06 $0.11 52m59s 25m22s
Finance Agent (v2) $1.66 $0.07 16m41s 7m02s
Tax Agent Bench $0.98 $0.05 10m58s 17m19s
MedCode N/A N/A 2m07s 93.37s
MedScribe N/A N/A 70.94s 74.48s
ProofBench v1.1 $0.76 $0.12 14m22s 1h03m
BioMysteryBench $1.23 $0.05 14m18s 28m28s
MysteryMechanism $1.31 $0.06 27m55s 35m08s
Terminal-Bench Science $4.82 $0.24 1h18m 3h59m
SAGE N/A N/A 3m22s 2m16s
Code Migration $14.58 $0.49 1h30m 3h12m
IOI $8.18 $0.22 3h57m 1h49m
ProgramBench $16.76 $0.47 4h27m 3h22m
Terminal-Bench 4.0 $5.07 $0.22 29m27s 2h25m
Vibe Code Bench v1.1 $4.88 $0.56 25m30s 50m13s
CyberBench v1.1 $4.07 $0.05 36m52s 28m22s
Public Benefits Bench v1.1 $0.92 $0.03 25m53s 20m56s

Results available only for Grok 4.6

  • Vals RSI Index
  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
  • Time Horizon Index: KSP

Results available only for MiMo V2.6 Flash

  • SRE Bench
Model details Grok 4.6 Model details MiMo V2.6 Flash