MiMo V2.5 vs GLM 5.3 Flash: Benchmark Comparison

MiMo V2.5 has the higher score on 2 of 16 shared benchmarks; GLM 5.3 Flash leads on 14.

The largest observed score gap is 36.06 pts on Legal Research Bench , where GLM 5.3 Flash leads.

Reported ±1 standard-error ranges overlap on 4 of 16 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark MiMo V2.5 GLM 5.3 Flash Gap Reported uncertainty
Harvey's Legal Agent Benchmark 1.67% ±0.00 6.67% ±1.65 5.00 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 9.13% ±2.00 45.19% ±3.46 36.06 pts Reported ±1 SE ranges do not overlap
LegalBench 78.89% ±0.51 83.93% ±0.43 5.04 pts Reported ±1 SE ranges do not overlap
EMB 55.09% ±3.13 55.93% ±3.56 0.84 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 36.73% ±0.11 57.85% ±2.06 21.12 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 29.16% ±2.71 61.96% ±3.26 32.79 pts Reported ±1 SE ranges do not overlap
TaxEval v2 71.83% ±0.88 75.59% ±0.84 3.76 pts Reported ±1 SE ranges do not overlap
MedScribe 72.15% ±1.85 88.94% ±1.91 16.79 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 16.00% ±3.69 21.00% ±4.09 5.00 pts Reported ±1 SE ranges overlap
GPQA Diamond 81.57% ±2.05 86.36% ±2.13 4.80 pts Reported ±1 SE ranges do not overlap
MMLU Pro 82.93% ±0.37 86.06% ±0.34 3.13 pts Reported ±1 SE ranges do not overlap
MMMU Pro 80.00% ±0.96 86.01% ±0.83 6.01 pts Reported ±1 SE ranges do not overlap
Code Migration 14.23% ±3.69 20.52% ±4.21 6.29 pts Reported ±1 SE ranges overlap
LiveCodeBench 81.51% ±1.07 80.51% ±1.05 1.00 pts Reported ±1 SE ranges overlap
SWE-bench 71.00% ±2.03 92.00% ±1.21 21.00 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 42.17% ±4.57 30.76% ±5.28 11.41 pts Reported ±1 SE ranges do not overlap

Performance by category

Category MiMo V2.5 average GLM 5.3 Flash average
Legal 29.90% 45.26%
Finance 48.20% 62.83%
Healthcare 72.15% 88.94%
Math 16.00% 21.00%
Academic 81.50% 86.14%
Coding 52.23% 55.95%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark MiMo V2.5 cost GLM 5.3 Flash cost MiMo V2.5 latency GLM 5.3 Flash latency
Harvey's Legal Agent Benchmark $0.04 $0.57 6m59s 52m17s
Legal Research Bench $0.03 $0.12 6m31s 51m07s
LegalBench N/A N/A 11.18s 12.52s
EMB $0.08 $0.71 30m38s 1h15m
Finance Agent (v2) $0.09 $0.05 5m39s 14m51s
Tax Agent Bench $0.02 $0.14 8m21s 21m37s
TaxEval v2 N/A N/A 26.30s 53.06s
MedScribe N/A N/A 20.26s 86.98s
ProofBench v1.1 $0.06 $0.22 20m56s 1h11m
GPQA Diamond N/A N/A 89.39s 2m11s
MMLU Pro N/A N/A 30.25s 44.76s
MMMU Pro N/A N/A 42.47s 28.22s
Code Migration $0.10 $3.13 37m55s 3h26m
LiveCodeBench N/A N/A 108.61s 5m02s
SWE-bench $0.01 $0.02 4m04s 1h04m
Vibe Code Bench v1.1 $0.07 $2.35 26m20s 1h24m

Results available only for MiMo V2.5

  • MortgageTax
  • MedCode
  • SAGE

Results available only for GLM 5.3 Flash

  • Terminal-Bench Science
  • IOI
  • ProgramBench
  • SkillsBench
  • Vibe Code Bench 1-100
  • CyberBench v1.1
Model details MiMo V2.5 Model details GLM 5.3 Flash