Claude Fable 5.1 vs MiMo V2.6 Flash: Benchmark Comparison

Claude Fable 5.1 has the higher score on 18 of 20 shared benchmarks; MiMo V2.6 Flash leads on 2.

The largest observed score gap is 43.11 pts on IOI , where Claude Fable 5.1 leads.

Reported ±1 standard-error ranges overlap on 2 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Fable 5.1 MiMo V2.6 Flash Gap Reported uncertainty
Vals Index 65.83% ±1.11 53.23% ±1.10 12.60 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 6.67% ±1.17 11.25% ±2.54 4.58 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 55.29% ±3.46 37.98% ±3.37 17.31 pts Reported ±1 SE ranges do not overlap
EMB 76.67% ±2.08 65.46% ±2.71 11.21 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 58.88% ±2.06 56.28% ±0.46 2.60 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 77.64% ±2.83 59.90% ±3.13 17.74 pts Reported ±1 SE ranges do not overlap
MedCode 53.51% ±2.17 41.06% ±2.03 12.45 pts Reported ±1 SE ranges do not overlap
MedScribe 91.29% ±1.95 85.28% ±1.99 6.02 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 100.00% ±0.00 63.00% ±4.85 37.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 47.75% ±3.36 21.62% ±2.77 26.13 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 40.00% ±5.90 4.29% ±2.44 35.71 pts Reported ±1 SE ranges do not overlap
SAGE 48.53% ±3.34 43.53% ±3.43 5.00 pts Reported ±1 SE ranges overlap
Code Migration 54.61% ±4.81 40.93% ±4.35 13.68 pts Reported ±1 SE ranges do not overlap
IOI 90.78% ±4.65 47.67% ±4.79 43.11 pts Reported ±1 SE ranges do not overlap
ProgramBench 7.00% ±1.81 0.50% ±0.50 6.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 58.08% ±3.31 24.24% ±1.51 33.84 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 90.26% ±1.57 78.96% ±4.04 11.30 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 70.42% ±5.43 75.36% ±5.42 4.94 pts Reported ±1 SE ranges overlap
SRE Bench 22.90% ±2.60 4.20% ±1.24 18.70 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 74.90% ±1.13 67.59% ±1.22 7.31 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Fable 5.1 average MiMo V2.6 Flash average
Legal 30.98% 24.62%
Finance 71.06% 60.55%
Healthcare 72.40% 63.17%
Math 100.00% 63.00%
Science 43.87% 12.95%
Education 48.53% 43.53%
Coding 60.15% 38.46%
Cyber 46.66% 39.78%
Social Mobility 74.90% 67.59%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Fable 5.1 cost MiMo V2.6 Flash cost Claude Fable 5.1 latency MiMo V2.6 Flash latency
Vals Index $28.71 $0.20 1h17m 58m02s
Harvey's Legal Agent Benchmark $46.21 $0.09 1h52m 17m12s
Legal Research Bench $23.06 $0.07 57m29s 15m20s
EMB $15.93 $0.11 29m32s 25m22s
Finance Agent (v2) $8.35 $0.07 18m14s 7m02s
Tax Agent Bench $13.18 $0.05 28m34s 17m19s
MedCode N/A N/A 3m35s 93.37s
MedScribe N/A N/A 3m08s 74.48s
ProofBench v1.1 $2.70 $0.12 7m58s 1h03m
MysteryMechanism $5.63 $0.06 17m41s 35m08s
Terminal-Bench Science $38.01 $0.24 2h36m 3h59m
SAGE N/A N/A 105.75s 2m16s
Code Migration $70.97 $0.49 4h20m 3h12m
IOI $11.20 $0.22 30m06s 1h49m
ProgramBench $58.52 $0.47 2h27m 3h22m
Terminal-Bench 4.0 $17.18 $0.22 1h00m 2h25m
Vibe Code Bench v1.1 $33.37 $0.56 57m40s 50m13s
CyberBench v1.1 $4.01 $0.05 18m32s 28m22s
SRE Bench $32.23 $0.40 2h03m 3h22m
Public Benefits Bench v1.1 $7.48 $0.03 48m02s 20m56s

Results available only for Claude Fable 5.1

  • Vals RSI Index
  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • Vibe Code Bench 1-100
  • CUA-bench
  • Time Horizon Index: KSP

Results available only for MiMo V2.6 Flash

  • BioMysteryBench
Model details Claude Fable 5.1 Model details MiMo V2.6 Flash