Model comparison

Gemini 3.7 Flash vs MiMo V2.6 Pro: Benchmark Comparison

Gemini 3.7 Flash has the higher score on 7 of 18 shared benchmarks; MiMo V2.6 Pro leads on 11.

The largest observed score gap is 29.11 pts on CyberBench v1.1 , where MiMo V2.6 Pro leads.

Reported ±1 standard-error ranges overlap on 7 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.7 Flash MiMo V2.6 Pro Gap Reported uncertainty
Vals Index 59.31% ±1.06 59.47% ±1.18 0.16 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 8.75% ±2.15 10.83% ±2.17 2.08 pts Reported ±1 SE ranges overlap
Legal Research Bench 34.62% ±3.31 47.12% ±3.47 12.50 pts Reported ±1 SE ranges do not overlap
EMB 71.33% ±2.26 62.86% ±3.14 8.47 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 59.04% ±0.27 57.34% ±0.57 1.70 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 57.66% ±3.31 64.94% ±3.21 7.27 pts Reported ±1 SE ranges do not overlap
MedCode 53.39% ±2.12 44.97% ±2.10 8.42 pts Reported ±1 SE ranges do not overlap
MedScribe 83.94% ±2.00 88.31% ±1.94 4.37 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 58.00% ±4.96 70.00% ±4.61 12.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 5.71% ±2.79 2.86% ±2.01 2.86 pts Reported ±1 SE ranges overlap
SAGE 49.23% ±3.38 45.05% ±3.40 4.18 pts Reported ±1 SE ranges overlap
Code Migration 34.80% ±4.22 43.01% ±4.32 8.21 pts Reported ±1 SE ranges overlap
IOI 67.83% ±3.61 39.33% ±2.41 28.50 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 6.06% ±0.88 24.75% ±3.07 18.69 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 70.39% ±4.84 85.22% ±3.39 14.83 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 43.75% ±2.21 72.86% ±5.50 29.11 pts Reported ±1 SE ranges do not overlap
SRE Bench 4.58% ±1.29 3.05% ±1.06 1.53 pts Reported ±1 SE ranges overlap

Performance by category

Category Gemini 3.7 Flash average MiMo V2.6 Pro average
Index 59.31% 59.47%
Legal 21.68% 28.97%
Finance 62.68% 61.71%
Healthcare 68.67% 66.64%
Math 58.00% 70.00%
Science 5.71% 2.86%
Education 49.23% 45.05%
Coding 35.82% 38.56%
Beta 24.16% 37.95%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.7 Flash cost MiMo V2.6 Pro cost Gemini 3.7 Flash latency MiMo V2.6 Pro latency
Vals Index $4.17 $0.39 43m13s 52m20s
Harvey's Legal Agent Benchmark $2.56 $0.22 7m32s 25m49s
Legal Research Bench $0.76 $0.18 4m49s 30m16s
EMB $5.69 $0.33 10m51s 55m14s
Finance Agent (v2) $1.48 $0.20 3m01s 10m48s
Tax Agent Bench $0.48 $0.13 87.24s 25m48s
MedCode N/A N/A 9.16s 2m57s
MedScribe N/A N/A 16.82s 2m53s
ProofBench v1.1 $0.56 $0.25 7m04s 1h02m
Terminal-Bench Science $13.22 $0.68 1h08m 3h40m
SAGE N/A N/A 9.79s 3m40s
Code Migration $21.46 $0.69 2h39m 2h49m
IOI $3.40 $0.84 21m09s 2h26m
ProgramBench $7.13 $0.73 21m30s 3h00m
Terminal-Bench 4.0 $14.07 $0.50 1h59m 2h36m
Vibe Code Bench v1.1 $4.83 $1.04 23m43s 1h01m
CyberBench v1.1 $2.15 $0.09 7m34s 27m41s
SRE Bench $21.33 $0.53 1h50m 2h23m

Results available only for Gemini 3.7 Flash

  • Vals RSI Index
  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench

Results available only for MiMo V2.6 Pro

  • MysteryMechanism
  • Public Benefits Bench v1.1
Model details Gemini 3.7 Flash Model details MiMo V2.6 Pro