Qwen 3.7 Plus vs MiMo V2.5: Benchmark Comparison

Qwen 3.7 Plus has the higher score on 5 of 9 shared benchmarks; MiMo V2.5 leads on 4.

The largest observed score gap is 9.55 pts on Tax Agent Bench , where Qwen 3.7 Plus leads.

Reported ±1 standard-error ranges overlap on 4 of 9 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Qwen 3.7 Plus MiMo V2.5 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 0.00% ±0.00 1.67% ±0.00 1.67 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 16.35% ±2.57 9.13% ±2.00 7.21 pts Reported ±1 SE ranges do not overlap
EMB 49.34% ±3.18 55.09% ±3.13 5.75 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 38.22% ±1.04 36.73% ±0.11 1.49 pts Reported ±1 SE ranges do not overlap
MortgageTax 66.18% ±0.93 59.26% ±0.98 6.92 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 38.71% ±2.81 29.16% ±2.71 9.55 pts Reported ±1 SE ranges do not overlap
SAGE 39.25% ±3.38 43.27% ±3.39 4.01 pts Reported ±1 SE ranges overlap
Code Migration 12.86% ±2.93 14.23% ±3.69 1.37 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 46.39% ±4.61 42.17% ±4.57 4.22 pts Reported ±1 SE ranges overlap

Performance by category

Category Qwen 3.7 Plus average MiMo V2.5 average
Legal 8.17% 5.40%
Finance 48.11% 45.06%
Education 39.25% 43.27%
Coding 29.62% 28.20%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Qwen 3.7 Plus cost MiMo V2.5 cost Qwen 3.7 Plus latency MiMo V2.5 latency
Harvey's Legal Agent Benchmark $0.23 $0.04 8m49s 6m59s
Legal Research Bench $0.30 $0.03 13m52s 6m31s
EMB $0.60 $0.08 36m49s 30m38s
Finance Agent (v2) $0.36 $0.09 9m02s 5m39s
MortgageTax N/A N/A 60.55s 32.94s
Tax Agent Bench $0.09 $0.02 5m10s 8m21s
SAGE N/A N/A 2m57s 52.66s
Code Migration $0.42 $0.10 35m47s 37m55s
Vibe Code Bench v1.1 $1.08 $0.07 37m12s 26m20s

Results available only for Qwen 3.7 Plus

  • SkillsBench

Results available only for MiMo V2.5

  • LegalBench
  • TaxEval v2
  • MedCode
  • MedScribe
  • ProofBench v1.1
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SWE-bench
Model details Qwen 3.7 Plus Model details MiMo V2.5