MiniMax-M3 vs MiMo V2.5: Benchmark Comparison

MiniMax-M3 has the higher score on 18 of 19 shared benchmarks; MiMo V2.5 leads on 1.

The largest observed score gap is 20.67 pts on Legal Research Bench , where MiniMax-M3 leads.

Reported ±1 standard-error ranges overlap on 6 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark MiniMax-M3 MiMo V2.5 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 4.17% ±1.65 1.67% ±0.00 2.50 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 29.81% ±3.18 9.13% ±2.00 20.67 pts Reported ±1 SE ranges do not overlap
LegalBench 85.42% ±0.44 78.89% ±0.51 6.53 pts Reported ±1 SE ranges do not overlap
EMB 47.76% ±2.95 55.09% ±3.13 7.33 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 48.27% ±0.44 36.73% ±0.11 11.54 pts Reported ±1 SE ranges do not overlap
MortgageTax 68.36% ±0.91 59.26% ±0.98 9.10 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 49.69% ±3.25 29.16% ±2.71 20.53 pts Reported ±1 SE ranges do not overlap
TaxEval v2 72.73% ±0.86 71.83% ±0.88 0.90 pts Reported ±1 SE ranges overlap
MedCode 46.29% ±2.10 31.89% ±2.02 14.39 pts Reported ±1 SE ranges do not overlap
MedScribe 87.25% ±1.96 72.15% ±1.85 15.10 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 18.00% ±3.86 16.00% ±3.69 2.00 pts Reported ±1 SE ranges overlap
GPQA Diamond 92.68% ±1.44 81.57% ±2.05 11.11 pts Reported ±1 SE ranges do not overlap
MMLU Pro 84.22% ±0.36 82.93% ±0.37 1.29 pts Reported ±1 SE ranges do not overlap
MMMU Pro 81.16% ±0.94 80.00% ±0.96 1.16 pts Reported ±1 SE ranges overlap
SAGE 50.57% ±3.44 43.27% ±3.39 7.31 pts Reported ±1 SE ranges do not overlap
Code Migration 19.93% ±3.94 14.23% ±3.69 5.70 pts Reported ±1 SE ranges overlap
LiveCodeBench 82.15% ±1.05 81.51% ±1.07 0.64 pts Reported ±1 SE ranges overlap
SWE-bench 75.00% ±1.94 71.00% ±2.03 4.00 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 47.57% ±5.44 42.17% ±4.57 5.39 pts Reported ±1 SE ranges overlap

Performance by category

Category MiniMax-M3 average MiMo V2.5 average
Legal 39.80% 29.90%
Finance 57.36% 50.41%
Healthcare 66.77% 52.02%
Math 18.00% 16.00%
Academic 86.02% 81.50%
Education 50.57% 43.27%
Coding 56.16% 52.23%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark MiniMax-M3 cost MiMo V2.5 cost MiniMax-M3 latency MiMo V2.5 latency
Harvey's Legal Agent Benchmark $1.46 $0.04 22m31s 6m59s
Legal Research Bench $0.34 $0.03 13m34s 6m31s
LegalBench N/A N/A 8.21s 11.18s
EMB $2.09 $0.08 31m30s 30m38s
Finance Agent (v2) $0.32 $0.09 8m17s 5m39s
MortgageTax N/A N/A 26.76s 32.94s
Tax Agent Bench $0.16 $0.02 5m15s 8m21s
TaxEval v2 N/A N/A 94.04s 26.30s
MedCode N/A N/A 63.12s 17.70s
MedScribe N/A N/A 2m04s 20.26s
ProofBench v1.1 $0.42 $0.06 11m59s 20m56s
GPQA Diamond N/A N/A 4m39s 89.39s
MMLU Pro N/A N/A 41.39s 30.25s
MMMU Pro N/A N/A 68.35s 42.47s
SAGE N/A N/A 118.54s 52.66s
Code Migration $7.07 $0.10 1h14m 37m55s
LiveCodeBench N/A N/A 6m07s 108.61s
SWE-bench $0.42 $0.01 12m06s 4m04s
Vibe Code Bench v1.1 $6.45 $0.07 1h25m 26m20s

Results available only for MiniMax-M3

  • Vals Index
  • SkillsBench
  • Terminal-Bench 4.0
  • Vibe Code Bench 1-100
  • CyberBench v1.1
  • Public Benefits Bench v1.1

Results available only for MiMo V2.5

None.

Model details MiniMax-M3 Model details MiMo V2.5