MiMo V2.5 Pro vs GLM 5.3 Flash: Benchmark Comparison

MiMo V2.5 Pro has the higher score on 4 of 15 shared benchmarks; GLM 5.3 Flash leads on 11.

The largest observed score gap is 29.33 pts on Legal Research Bench , where GLM 5.3 Flash leads.

Reported ±1 standard-error ranges overlap on 6 of 15 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark MiMo V2.5 Pro GLM 5.3 Flash Gap Reported uncertainty
Harvey's Legal Agent Benchmark 2.08% ±0.83 6.67% ±1.65 4.58 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 15.87% ±2.54 45.19% ±3.46 29.33 pts Reported ±1 SE ranges do not overlap
LegalBench 77.13% ±0.55 83.93% ±0.43 6.80 pts Reported ±1 SE ranges do not overlap
EMB 55.23% ±3.11 55.93% ±3.56 0.70 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 41.50% ±1.19 57.85% ±2.06 16.35 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 34.79% ±2.62 61.96% ±3.26 27.17 pts Reported ±1 SE ranges do not overlap
TaxEval v2 73.79% ±0.86 75.59% ±0.84 1.80 pts Reported ±1 SE ranges do not overlap
MedScribe 83.73% ±2.06 88.94% ±1.91 5.21 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 22.00% ±4.16 21.00% ±4.09 1.00 pts Reported ±1 SE ranges overlap
GPQA Diamond 82.58% ±2.47 86.36% ±2.13 3.79 pts Reported ±1 SE ranges overlap
MMLU Pro 84.59% ±0.45 86.06% ±0.34 1.46 pts Reported ±1 SE ranges do not overlap
Code Migration 21.56% ±4.13 20.52% ±4.21 1.04 pts Reported ±1 SE ranges overlap
LiveCodeBench 81.35% ±1.07 80.51% ±1.05 0.84 pts Reported ±1 SE ranges overlap
SWE-bench 74.00% ±1.96 92.00% ±1.21 18.00 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 34.11% ±4.53 30.76% ±5.28 3.35 pts Reported ±1 SE ranges overlap

Performance by category

Category MiMo V2.5 Pro average GLM 5.3 Flash average
Legal 31.69% 45.26%
Finance 51.33% 62.83%
Healthcare 83.73% 88.94%
Math 22.00% 21.00%
Academic 83.59% 86.21%
Coding 52.76% 55.95%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark MiMo V2.5 Pro cost GLM 5.3 Flash cost MiMo V2.5 Pro latency GLM 5.3 Flash latency
Harvey's Legal Agent Benchmark $0.11 $0.57 9m04s 52m17s
Legal Research Bench $0.08 $0.12 6m56s 51m07s
LegalBench N/A N/A 27.14s 12.52s
EMB $0.22 $0.71 29m08s 1h15m
Finance Agent (v2) $0.21 $0.05 8m17s 14m51s
Tax Agent Bench $0.04 $0.14 6m17s 21m37s
TaxEval v2 N/A N/A 62.83s 53.06s
MedScribe N/A N/A 90.71s 86.98s
ProofBench v1.1 $0.09 $0.22 19m37s 1h11m
GPQA Diamond N/A N/A 4m11s 2m11s
MMLU Pro N/A N/A 67.76s 44.76s
Code Migration $0.24 $3.13 48m53s 3h26m
LiveCodeBench N/A N/A 4m31s 5m02s
SWE-bench $0.02 $0.02 7m51s 1h04m
Vibe Code Bench v1.1 $0.17 $2.35 40m04s 1h24m

Results available only for MiMo V2.5 Pro

  • Vals Index
  • MedCode
  • Terminal-Bench 4.0

Results available only for GLM 5.3 Flash

  • Terminal-Bench Science
  • MMMU Pro
  • IOI
  • ProgramBench
  • SkillsBench
  • Vibe Code Bench 1-100
  • CyberBench v1.1
Model details MiMo V2.5 Pro Model details GLM 5.3 Flash