Model comparison

Claude Sonnet 5 vs MiMo V2.6 Pro: Benchmark Comparison

Claude Sonnet 5 has the higher score on 6 of 18 shared benchmarks; MiMo V2.6 Pro leads on 11.

The largest observed score gap is 21.72 pts on Terminal-Bench 4.0 , where MiMo V2.6 Pro leads.

Reported ±1 standard-error ranges overlap on 11 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5 MiMo V2.6 Pro Gap Reported uncertainty
Vals Index 51.77% ±1.09 55.20% ±1.18 3.42 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 5.00% ±1.65 10.83% ±2.17 5.83 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.83% ±3.43 47.12% ±3.47 5.29 pts Reported ±1 SE ranges overlap
EMB 66.32% ±3.01 62.86% ±3.14 3.46 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.91% ±0.52 57.34% ±0.57 3.43 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.27% ±3.19 64.94% ±3.21 2.67 pts Reported ±1 SE ranges overlap
MedCode 47.54% ±2.27 44.97% ±2.10 2.57 pts Reported ±1 SE ranges overlap
MedScribe 76.05% ±3.05 88.31% ±1.94 12.25 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 77.00% ±4.23 70.00% ±4.61 7.00 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 2.86% ±2.01 2.86% ±2.01 0.00 pts Reported ±1 SE ranges overlap
SAGE 48.92% ±3.40 45.05% ±3.40 3.88 pts Reported ±1 SE ranges overlap
Code Migration 44.39% ±4.25 43.01% ±4.32 1.38 pts Reported ±1 SE ranges overlap
IOI 45.00% ±2.75 39.33% ±2.41 5.67 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 9.60% ±1.01 31.31% ±3.07 21.72 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 81.33% ±3.05 85.22% ±3.39 3.90 pts Reported ±1 SE ranges overlap
CyberBench v1.1 61.91% ±5.74 72.86% ±5.50 10.95 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 66.03% ±1.23 68.94% ±1.20 2.91 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Sonnet 5 average MiMo V2.6 Pro average
Legal 23.41% 28.97%
Finance 60.83% 61.71%
Healthcare 61.80% 66.64%
Math 77.00% 70.00%
Science 2.86% 2.86%
Education 48.92% 45.05%
Coding 36.06% 39.88%
Beta 61.91% 72.86%
Social Mobility 66.03% 68.94%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5 cost MiMo V2.6 Pro cost Claude Sonnet 5 latency MiMo V2.6 Pro latency
Vals Index $13.72 $0.41 54m16s 1h07m
Harvey's Legal Agent Benchmark $8.95 $0.22 38m52s 25m49s
Legal Research Bench $2.72 $0.18 25m46s 30m16s
EMB $10.29 $0.33 49m46s 55m14s
Finance Agent (v2) $0.75 $0.20 13m12s 10m48s
Tax Agent Bench $1.79 $0.13 19m19s 25m48s
MedCode N/A N/A 2m15s 2m57s
MedScribe N/A N/A 4m12s 2m53s
ProofBench v1.1 $1.37 $0.25 14m58s 1h02m
Terminal-Bench Science $18.70 $0.68 2h55m 3h40m
SAGE N/A N/A 7m14s 3m40s
Code Migration $35.31 $0.69 1h57m 2h49m
IOI $12.89 $0.84 59m44s 2h26m
ProgramBench $24.36 $0.73 1h31m 3h00m
Terminal-Bench 4.0 $26.33 $0.50 1h45m 2h42m
Vibe Code Bench v1.1 $25.39 $1.04 1h07m 1h01m
CyberBench v1.1 $1.89 $0.09 16m57s 27m41s
Public Benefits Bench v1.1 $1.29 $0.08 21m09s 37m14s

Results available only for Claude Sonnet 5

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100

Results available only for MiMo V2.6 Pro

  • MysteryMechanism
  • SRE Bench
Model details Claude Sonnet 5 Model details MiMo V2.6 Pro