Claude Fable 5 vs MiMo V2.6 Flash: Benchmark Comparison

Claude Fable 5 has the higher score on 15 of 16 shared benchmarks; MiMo V2.6 Flash leads on 0.

The largest observed score gap is 32.00 pts on ProofBench v1.1 , where Claude Fable 5 leads.

Reported ±1 standard-error ranges overlap on 3 of 16 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Fable 5 MiMo V2.6 Flash Gap Reported uncertainty
Vals Index 61.39% ±1.00 53.23% ±1.10 8.16 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 11.25% ±2.15 11.25% ±2.54 0.00 pts Reported ±1 SE ranges overlap
Legal Research Bench 49.52% ±3.48 37.98% ±3.37 11.54 pts Reported ±1 SE ranges do not overlap
EMB 73.67% ±2.47 65.46% ±2.71 8.21 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 56.31% ±0.84 56.28% ±0.46 0.04 pts Reported ±1 SE ranges overlap
Tax Agent Bench 69.82% ±3.12 59.90% ±3.13 9.92 pts Reported ±1 SE ranges do not overlap
MedCode 56.07% ±2.20 41.06% ±2.03 15.01 pts Reported ±1 SE ranges do not overlap
MedScribe 88.52% ±1.95 85.28% ±1.99 3.25 pts Reported ±1 SE ranges overlap
ProofBench v1.1 95.00% ±2.19 63.00% ±4.85 32.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 15.71% ±4.38 4.29% ±2.44 11.43 pts Reported ±1 SE ranges do not overlap
SAGE 51.89% ±3.40 43.53% ±3.43 8.36 pts Reported ±1 SE ranges do not overlap
Code Migration 55.06% ±4.61 40.93% ±4.35 14.14 pts Reported ±1 SE ranges do not overlap
ProgramBench 2.00% ±0.99 0.50% ±0.50 1.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 41.41% ±1.34 24.24% ±1.51 17.17 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 90.35% ±2.10 78.96% ±4.04 11.39 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 70.43% ±1.19 67.59% ±1.22 2.84 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Fable 5 average MiMo V2.6 Flash average
Legal 30.38% 24.62%
Finance 66.60% 60.55%
Healthcare 72.30% 63.17%
Math 95.00% 63.00%
Science 15.71% 4.29%
Education 51.89% 43.53%
Coding 47.21% 36.16%
Social Mobility 70.43% 67.59%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Fable 5 cost MiMo V2.6 Flash cost Claude Fable 5 latency MiMo V2.6 Flash latency
Vals Index $29.59 $0.20 42m24s 58m02s
Harvey's Legal Agent Benchmark $19.23 $0.09 26m53s 17m12s
Legal Research Bench $9.79 $0.07 22m18s 15m20s
EMB $12.35 $0.11 26m34s 25m22s
Finance Agent (v2) $8.06 $0.07 10m11s 7m02s
Tax Agent Bench $6.70 $0.05 14m27s 17m19s
MedCode N/A N/A 91.44s 93.37s
MedScribe N/A N/A 119.47s 74.48s
ProofBench v1.1 $5.45 $0.12 13m43s 1h03m
Terminal-Bench Science $51.99 $0.24 1h59m 3h59m
SAGE N/A N/A 116.92s 2m16s
Code Migration $112.10 $0.49 1h49m 3h12m
ProgramBench $75.68 $0.47 2h37m 3h22m
Terminal-Bench 4.0 $30.34 $0.22 1h08m 2h25m
Vibe Code Bench v1.1 $41.71 $0.56 1h02m 50m13s
Public Benefits Bench v1.1 $4.59 $0.03 21m17s 20m56s

Results available only for Claude Fable 5

  • Vals RSI Index
  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SWE-bench

Results available only for MiMo V2.6 Flash

  • BioMysteryBench
  • MysteryMechanism
  • IOI
  • CyberBench v1.1
  • SRE Bench
Model details Claude Fable 5 Model details MiMo V2.6 Flash