Gemini 3.1 Pro Preview (02/26) vs MiMo V2.5: Benchmark Comparison

Gemini 3.1 Pro Preview (02/26) has the higher score on 16 of 19 shared benchmarks; MiMo V2.5 leads on 3.

The largest observed score gap is 27.17 pts on MedCode , where Gemini 3.1 Pro Preview (02/26) leads.

Reported ±1 standard-error ranges overlap on 4 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.1 Pro Preview (02/26) MiMo V2.5 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 0.00% ±0.00 1.67% ±0.00 1.67 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 20.67% ±2.81 9.13% ±2.00 11.54 pts Reported ±1 SE ranges do not overlap
LegalBench 87.40% ±0.33 78.89% ±0.51 8.51 pts Reported ±1 SE ranges do not overlap
EMB 52.62% ±2.97 55.09% ±3.13 2.48 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 42.98% ±1.21 36.73% ±0.11 6.25 pts Reported ±1 SE ranges do not overlap
MortgageTax 69.40% ±0.91 59.26% ±0.98 10.14 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 37.84% ±3.10 29.16% ±2.71 8.68 pts Reported ±1 SE ranges do not overlap
TaxEval v2 72.88% ±0.86 71.83% ±0.88 1.05 pts Reported ±1 SE ranges overlap
MedCode 59.06% ±2.00 31.89% ±2.02 27.17 pts Reported ±1 SE ranges do not overlap
MedScribe 76.11% ±1.92 72.15% ±1.85 3.96 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 26.00% ±4.41 16.00% ±3.69 10.00 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 95.45% ±1.05 81.57% ±2.05 13.89 pts Reported ±1 SE ranges do not overlap
MMLU Pro 90.99% ±0.28 82.93% ±0.37 8.06 pts Reported ±1 SE ranges do not overlap
MMMU Pro 88.21% ±0.78 80.00% ±0.96 8.21 pts Reported ±1 SE ranges do not overlap
SAGE 48.68% ±3.29 43.27% ±3.39 5.41 pts Reported ±1 SE ranges overlap
Code Migration 17.31% ±3.93 14.23% ±3.69 3.08 pts Reported ±1 SE ranges overlap
LiveCodeBench 88.48% ±0.93 81.51% ±1.07 6.97 pts Reported ±1 SE ranges do not overlap
SWE-bench 78.80% ±1.83 71.00% ±2.03 7.80 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 32.03% ±4.34 42.17% ±4.57 10.14 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Gemini 3.1 Pro Preview (02/26) average MiMo V2.5 average
Legal 36.02% 29.90%
Finance 55.14% 50.41%
Healthcare 67.59% 52.02%
Math 26.00% 16.00%
Academic 91.55% 81.50%
Education 48.68% 43.27%
Coding 54.16% 52.23%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.1 Pro Preview (02/26) cost MiMo V2.5 cost Gemini 3.1 Pro Preview (02/26) latency MiMo V2.5 latency
Harvey's Legal Agent Benchmark $1.15 $0.04 7m17s 6m59s
Legal Research Bench $1.10 $0.03 4m02s 6m31s
LegalBench N/A N/A 10.06s 11.18s
EMB $3.77 $0.08 11m49s 30m38s
Finance Agent (v2) $1.59 $0.09 3m28s 5m39s
MortgageTax N/A N/A 23.62s 32.94s
Tax Agent Bench $0.30 $0.02 100.96s 8m21s
TaxEval v2 N/A N/A 41.93s 26.30s
MedCode N/A N/A 38.52s 17.70s
MedScribe N/A N/A 69.14s 20.26s
ProofBench v1.1 $0.78 $0.06 14m49s 20m56s
GPQA Diamond N/A N/A 65.76s 89.39s
MMLU Pro N/A N/A 23.91s 30.25s
MMMU Pro N/A N/A 76.99s 42.47s
SAGE N/A N/A 61.17s 52.66s
Code Migration $1.61 $0.10 12m01s 37m55s
LiveCodeBench N/A N/A 88.89s 108.61s
SWE-bench $0.78 $0.01 5m12s 4m04s
Vibe Code Bench v1.1 $3.83 $0.07 20m12s 26m20s

Results available only for Gemini 3.1 Pro Preview (02/26)

  • Vals Index
  • BioMysteryBench
  • MysteryMechanism
  • Terminal-Bench Science
  • IOI
  • ProgramBench
  • Terminal-Bench 4.0
  • Vibe Code Bench 1-100
  • Public Benefits Bench v1.1

Results available only for MiMo V2.5

None.

Model details Gemini 3.1 Pro Preview (02/26) Model details MiMo V2.5