Claude Opus 5.5 vs MiMo V2.6 Flash: Benchmark Comparison

Claude Opus 5.5 has the higher score on 19 of 21 shared benchmarks; MiMo V2.6 Flash leads on 2.

The largest observed score gap is 47.39 pts on IOI , where Claude Opus 5.5 leads.

Reported ±1 standard-error ranges overlap on 1 of 21 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Opus 5.5 MiMo V2.6 Flash Gap Reported uncertainty
Vals Index 66.97% ±0.89 53.23% ±1.10 13.74 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 3.75% ±1.47 11.25% ±2.54 7.50 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 50.48% ±3.48 37.98% ±3.37 12.50 pts Reported ±1 SE ranges do not overlap
EMB 75.94% ±2.38 65.46% ±2.71 10.48 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 58.59% ±0.17 56.28% ±0.46 2.31 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 70.50% ±3.15 59.90% ±3.13 10.59 pts Reported ±1 SE ranges do not overlap
MedCode 49.80% ±2.27 41.06% ±2.03 8.74 pts Reported ±1 SE ranges do not overlap
MedScribe 91.43% ±1.93 85.28% ±1.99 6.16 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 100.00% ±0.00 63.00% ±4.85 37.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 79.26% ±0.98 69.26% ±0.98 10.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 49.55% ±3.36 21.62% ±2.77 27.93 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 47.14% ±6.01 4.29% ±2.44 42.86 pts Reported ±1 SE ranges do not overlap
SAGE 45.83% ±3.36 43.53% ±3.43 2.30 pts Reported ±1 SE ranges overlap
Code Migration 66.65% ±4.33 40.93% ±4.35 25.72 pts Reported ±1 SE ranges do not overlap
IOI 95.06% ±4.94 47.67% ±4.79 47.39 pts Reported ±1 SE ranges do not overlap
ProgramBench 18.50% ±2.75 0.50% ±0.50 18.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 65.15% ±0.00 24.24% ±1.51 40.91 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 90.29% ±1.53 78.96% ±4.04 11.33 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 55.36% ±5.13 75.36% ±5.42 20.00 pts Reported ±1 SE ranges do not overlap
SRE Bench 33.59% ±2.92 4.20% ±1.24 29.39 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 70.64% ±1.19 67.59% ±1.22 3.05 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Opus 5.5 average MiMo V2.6 Flash average
Legal 27.12% 24.62%
Finance 68.34% 60.55%
Healthcare 70.61% 63.17%
Math 100.00% 63.00%
Science 58.65% 31.72%
Education 45.83% 43.53%
Coding 67.13% 38.46%
Cyber 44.47% 39.78%
Social Mobility 70.64% 67.59%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Opus 5.5 cost MiMo V2.6 Flash cost Claude Opus 5.5 latency MiMo V2.6 Flash latency
Vals Index $32.14 $0.20 1h19m 58m02s
Harvey's Legal Agent Benchmark $21.38 $0.09 49m00s 17m12s
Legal Research Bench $25.16 $0.07 1h52m 15m20s
EMB $11.02 $0.11 35m06s 25m22s
Finance Agent (v2) $9.22 $0.07 33m08s 7m02s
Tax Agent Bench $15.09 $0.05 1h07m 17m19s
MedCode N/A N/A 4m07s 93.37s
MedScribe N/A N/A 7m22s 74.48s
ProofBench v1.1 $0.96 $0.12 5m51s 1h03m
BioMysteryBench $3.64 $0.05 12m19s 28m28s
MysteryMechanism $4.62 $0.06 19m38s 35m08s
Terminal-Bench Science $19.12 $0.24 2h28m 3h59m
SAGE N/A N/A 114.20s 2m16s
Code Migration $112.97 $0.49 3h19m 3h12m
IOI $5.26 $0.22 17m15s 1h49m
ProgramBench $68.05 $0.47 2h24m 3h22m
Terminal-Bench 4.0 $13.20 $0.22 1h04m 2h25m
Vibe Code Bench v1.1 $57.92 $0.56 1h36m 50m13s
CyberBench v1.1 $3.48 $0.05 18m11s 28m22s
SRE Bench $34.65 $0.40 1h37m 3h22m
Public Benefits Bench v1.1 $5.95 $0.03 1h20m 20m56s

Results available only for Claude Opus 5.5

  • Vals RSI Index
  • Vibe Code Bench 1-100
  • CUA-bench
  • Time Horizon Index: KSP

Results available only for MiMo V2.6 Flash

None.

Model details Claude Opus 5.5 Model details MiMo V2.6 Flash