MiniMax-M3 vs MiMo V2.5 Pro: Benchmark Comparison

MiniMax-M3 has the higher score on 13 of 18 shared benchmarks; MiMo V2.5 Pro leads on 5.

The largest observed score gap is 14.90 pts on Tax Agent Bench , where MiniMax-M3 leads.

Reported ±1 standard-error ranges overlap on 9 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark MiniMax-M3 MiMo V2.5 Pro Gap Reported uncertainty
Vals Index 36.53% ±1.17 33.84% ±1.17 2.69 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 4.17% ±1.65 2.08% ±0.83 2.08 pts Reported ±1 SE ranges overlap
Legal Research Bench 29.81% ±3.18 15.87% ±2.54 13.94 pts Reported ±1 SE ranges do not overlap
LegalBench 85.42% ±0.44 77.13% ±0.55 8.29 pts Reported ±1 SE ranges do not overlap
EMB 47.76% ±2.95 55.23% ±3.11 7.47 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 48.27% ±0.44 41.50% ±1.19 6.77 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 49.69% ±3.25 34.79% ±2.62 14.90 pts Reported ±1 SE ranges do not overlap
TaxEval v2 72.73% ±0.86 73.79% ±0.86 1.06 pts Reported ±1 SE ranges overlap
MedCode 46.29% ±2.10 32.48% ±1.91 13.80 pts Reported ±1 SE ranges do not overlap
MedScribe 87.25% ±1.96 83.73% ±2.06 3.52 pts Reported ±1 SE ranges overlap
ProofBench v1.1 18.00% ±3.86 22.00% ±4.16 4.00 pts Reported ±1 SE ranges overlap
GPQA Diamond 92.68% ±1.44 82.58% ±2.47 10.10 pts Reported ±1 SE ranges do not overlap
MMLU Pro 84.22% ±0.36 84.59% ±0.45 0.38 pts Reported ±1 SE ranges overlap
Code Migration 19.93% ±3.94 21.56% ±4.13 1.63 pts Reported ±1 SE ranges overlap
LiveCodeBench 82.15% ±1.05 81.35% ±1.07 0.80 pts Reported ±1 SE ranges overlap
SWE-bench 75.00% ±1.94 74.00% ±1.96 1.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 1.01% ±1.01 0.51% ±0.51 0.51 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 47.57% ±5.44 34.11% ±4.53 13.45 pts Reported ±1 SE ranges do not overlap

Performance by category

Category MiniMax-M3 average MiMo V2.5 Pro average
Legal 39.80% 31.69%
Finance 54.61% 51.33%
Healthcare 66.77% 58.11%
Math 18.00% 22.00%
Academic 88.45% 83.59%
Coding 45.13% 42.31%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark MiniMax-M3 cost MiMo V2.5 Pro cost MiniMax-M3 latency MiMo V2.5 Pro latency
Vals Index $3.08 $0.16 38m55s 34m33s
Harvey's Legal Agent Benchmark $1.46 $0.11 22m31s 9m04s
Legal Research Bench $0.34 $0.08 13m34s 6m56s
LegalBench N/A N/A 8.21s 27.14s
EMB $2.09 $0.22 31m30s 29m08s
Finance Agent (v2) $0.32 $0.21 8m17s 8m17s
Tax Agent Bench $0.16 $0.04 5m15s 6m17s
TaxEval v2 N/A N/A 94.04s 62.83s
MedCode N/A N/A 63.12s 38.45s
MedScribe N/A N/A 2m04s 90.71s
ProofBench v1.1 $0.42 $0.09 11m59s 19m37s
GPQA Diamond N/A N/A 4m39s 4m11s
MMLU Pro N/A N/A 41.39s 67.76s
Code Migration $7.07 $0.24 1h14m 48m53s
LiveCodeBench N/A N/A 6m07s 4m31s
SWE-bench $0.42 $0.02 12m06s 7m51s
Terminal-Bench 4.0 $7.23 $0.23 1h16m 2h07m
Vibe Code Bench v1.1 $6.45 $0.17 1h25m 40m04s

Results available only for MiniMax-M3

  • MortgageTax
  • MMMU Pro
  • SAGE
  • SkillsBench
  • Vibe Code Bench 1-100
  • CyberBench v1.1
  • Public Benefits Bench v1.1

Results available only for MiMo V2.5 Pro

None.

Model details MiniMax-M3 Model details MiMo V2.5 Pro