Claude Fable 5.1 vs Muse Spark 1.3 Max: Benchmark Comparison

Claude Fable 5.1 has the higher score on 14 of 18 shared benchmarks; Muse Spark 1.3 Max leads on 3.

The largest observed score gap is 42.00 pts on ProofBench v1.1 , where Claude Fable 5.1 leads.

Reported ±1 standard-error ranges overlap on 6 of 16 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Fable 5.1 Muse Spark 1.3 Max Gap Reported uncertainty
Vals Index 65.83% ±1.11 58.16% ±1.19 7.66 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 36.09% 19.64% 16.45 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 6.67% ±1.17 23.75% ±3.55 17.08 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 55.29% ±3.46 55.29% ±3.46 0.00 pts Reported ±1 SE ranges overlap
EMB 76.67% ±2.08 67.43% ±3.06 9.24 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 58.88% ±2.06 59.96% ±2.06 1.08 pts Reported ±1 SE ranges overlap
Tax Agent Bench 77.64% ±2.83 72.44% ±2.88 5.20 pts Reported ±1 SE ranges overlap
ProofBench v1.1 100.00% ±0.00 58.00% ±4.96 42.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 47.75% ±3.36 36.04% ±3.23 11.71 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 40.00% ±5.90 10.00% ±3.61 30.00 pts Reported ±1 SE ranges do not overlap
Code Migration 54.61% ±4.81 47.41% ±4.26 7.19 pts Reported ±1 SE ranges overlap
IOI 90.78% ±4.65 56.56% ±2.52 34.22 pts Reported ±1 SE ranges do not overlap
ProgramBench 7.00% ±1.81 2.50% ±1.11 4.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 58.08% ±3.31 24.75% ±0.51 33.33 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 28.00% ±4.49 20.46% ±3.85 7.54 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 90.26% ±1.57 85.86% ±2.51 4.41 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 70.42% ±5.43 72.74% ±5.67 2.32 pts Reported ±1 SE ranges overlap
CUA-bench 13.17% 5.83% 7.33 pts Uncertainty comparison unavailable

Performance by category

Category Claude Fable 5.1 average Muse Spark 1.3 Max average
Legal 30.98% 39.52%
Finance 71.06% 66.61%
Math 100.00% 58.00%
Science 43.87% 23.02%
Coding 54.79% 39.59%
Cyber 70.42% 72.74%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Fable 5.1 cost Muse Spark 1.3 Max cost Claude Fable 5.1 latency Muse Spark 1.3 Max latency
Vals Index $28.71 $3.79 1h17m 23m33s
Vals RSI Index $298.22 $48.82 90h00m 90h00m
Harvey's Legal Agent Benchmark $46.21 $2.26 1h52m 14m35s
Legal Research Bench $23.06 $0.59 57m29s 5m17s
EMB $15.93 $2.55 29m32s 13m50s
Finance Agent (v2) $8.35 $0.76 18m14s 3m19s
Tax Agent Bench $13.18 $0.39 28m34s 4m09s
ProofBench v1.1 $2.70 $0.48 7m58s 9m29s
MysteryMechanism $5.63 $0.85 17m41s 7m32s
Terminal-Bench Science $38.01 $6.09 2h36m 1h24m
Code Migration $70.97 $14.21 4h20m 1h24m
IOI $11.20 $1.73 30m06s 28m58s
ProgramBench $58.52 $54.29 2h27m 5h28m
Terminal-Bench 4.0 $17.18 $6.65 1h00m 52m56s
Vibe Code Bench 1-100 $149.66 $5.60 2h53m 20m02s
Vibe Code Bench v1.1 $33.37 $2.54 57m40s 16m18s
CyberBench v1.1 $4.01 $3.34 18m32s 16m18s
CUA-bench $323.78 $143.48 37.44s 36.39s

Results available only for Claude Fable 5.1

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • MedCode
  • MedScribe
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SAGE
  • LiveCodeBench
  • SkillsBench
  • SRE Bench
  • Public Benefits Bench v1.1
  • Time Horizon Index: KSP

Results available only for Muse Spark 1.3 Max

None.

Model details Claude Fable 5.1 Model details Muse Spark 1.3 Max