Model comparison

Claude Sonnet 5 vs MiMo V2.6 Flash: Benchmark Comparison

Claude Sonnet 5 has the higher score on 8 of 18 shared benchmarks; MiMo V2.6 Flash leads on 10.

The largest observed score gap is 14.65 pts on Terminal-Bench 4.0 , where MiMo V2.6 Flash leads.

Reported ±1 standard-error ranges overlap on 11 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5 MiMo V2.6 Flash Gap Reported uncertainty
Vals Index 51.77% ±1.09 53.23% ±1.10 1.46 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 5.00% ±1.65 11.25% ±2.29 6.25 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.83% ±3.43 37.98% ±3.37 3.85 pts Reported ±1 SE ranges overlap
EMB 66.32% ±3.01 65.46% ±2.71 0.86 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.91% ±0.52 56.28% ±0.46 2.37 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.27% ±3.19 59.90% ±3.13 2.37 pts Reported ±1 SE ranges overlap
MedCode 47.54% ±2.27 41.06% ±2.03 6.48 pts Reported ±1 SE ranges do not overlap
MedScribe 76.05% ±3.05 85.28% ±1.99 9.22 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 77.00% ±4.23 63.00% ±4.85 14.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 2.86% ±2.01 4.29% ±2.44 1.43 pts Reported ±1 SE ranges overlap
SAGE 48.92% ±3.40 43.53% ±3.43 5.39 pts Reported ±1 SE ranges overlap
Code Migration 44.39% ±4.25 40.93% ±4.35 3.47 pts Reported ±1 SE ranges overlap
IOI 45.00% ±2.75 47.67% ±4.79 2.67 pts Reported ±1 SE ranges overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 9.60% ±1.01 24.24% ±1.51 14.65 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 81.33% ±3.05 78.96% ±4.04 2.37 pts Reported ±1 SE ranges overlap
CyberBench v1.1 61.91% ±5.74 75.36% ±5.42 13.45 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 66.03% ±1.23 67.59% ±1.22 1.56 pts Reported ±1 SE ranges overlap

Performance by category

Category Claude Sonnet 5 average MiMo V2.6 Flash average
Legal 23.41% 24.62%
Finance 60.83% 60.55%
Healthcare 61.80% 63.17%
Math 77.00% 63.00%
Science 2.86% 4.29%
Education 48.92% 43.53%
Coding 36.06% 38.46%
Beta 61.91% 75.36%
Social Mobility 66.03% 67.59%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5 cost MiMo V2.6 Flash cost Claude Sonnet 5 latency MiMo V2.6 Flash latency
Vals Index $13.72 $0.20 54m16s 58m02s
Harvey's Legal Agent Benchmark $8.95 $0.09 38m52s 17m12s
Legal Research Bench $2.72 $0.07 25m46s 15m20s
EMB $10.29 $0.11 49m46s 25m22s
Finance Agent (v2) $0.75 $0.07 13m12s 7m02s
Tax Agent Bench $1.79 $0.05 19m19s 17m19s
MedCode N/A N/A 2m15s 93.37s
MedScribe N/A N/A 4m12s 74.48s
ProofBench v1.1 $1.37 $0.12 14m58s 1h03m
Terminal-Bench Science $18.70 $0.24 2h55m 3h59m
SAGE N/A N/A 7m14s 2m16s
Code Migration $35.31 $0.49 1h57m 3h12m
IOI $12.89 $0.22 59m44s 1h49m
ProgramBench $24.36 $0.47 1h31m 3h22m
Terminal-Bench 4.0 $26.33 $0.22 1h45m 2h25m
Vibe Code Bench v1.1 $25.39 $0.56 1h07m 50m13s
CyberBench v1.1 $1.89 $0.05 16m57s 28m22s
Public Benefits Bench v1.1 $1.29 $0.03 21m09s 20m56s

Results available only for Claude Sonnet 5

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100

Results available only for MiMo V2.6 Flash

  • BioMysteryBench
  • MysteryMechanism
  • SRE Bench
Model details Claude Sonnet 5 Model details MiMo V2.6 Flash