Claude Fable 5 vs Muse Spark 1.3 Max: Benchmark Comparison

Claude Fable 5 has the higher score on 8 of 13 shared benchmarks; Muse Spark 1.3 Max leads on 5.

The largest observed score gap is 37.00 pts on ProofBench v1.1 , where Claude Fable 5 leads.

Reported ±1 standard-error ranges overlap on 6 of 12 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Fable 5 Muse Spark 1.3 Max Gap Reported uncertainty
Vals Index 61.39% ±1.00 58.16% ±1.19 3.23 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 24.32% 19.64% 4.68 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 11.25% ±2.15 23.75% ±3.55 12.50 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 49.52% ±3.48 55.29% ±3.46 5.77 pts Reported ±1 SE ranges overlap
EMB 73.67% ±2.47 67.43% ±3.06 6.24 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 56.31% ±0.84 59.96% ±2.06 3.64 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 69.82% ±3.12 72.44% ±2.88 2.62 pts Reported ±1 SE ranges overlap
ProofBench v1.1 95.00% ±2.19 58.00% ±4.96 37.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 15.71% ±4.38 10.00% ±3.61 5.71 pts Reported ±1 SE ranges overlap
Code Migration 55.06% ±4.61 47.41% ±4.26 7.65 pts Reported ±1 SE ranges overlap
ProgramBench 2.00% ±0.99 2.50% ±1.11 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 41.41% ±1.34 24.75% ±0.51 16.67 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 90.35% ±2.10 85.86% ±2.51 4.50 pts Reported ±1 SE ranges overlap

Performance by category

Category Claude Fable 5 average Muse Spark 1.3 Max average
Legal 30.38% 39.52%
Finance 66.60% 66.61%
Math 95.00% 58.00%
Science 15.71% 10.00%
Coding 47.21% 40.13%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Fable 5 cost Muse Spark 1.3 Max cost Claude Fable 5 latency Muse Spark 1.3 Max latency
Vals Index $29.59 $3.79 42m24s 23m33s
Vals RSI Index $491.22 $48.82 90h00m 90h00m
Harvey's Legal Agent Benchmark $19.23 $2.26 26m53s 14m35s
Legal Research Bench $9.79 $0.59 22m18s 5m17s
EMB $12.35 $2.55 26m34s 13m50s
Finance Agent (v2) $8.06 $0.76 10m11s 3m19s
Tax Agent Bench $6.70 $0.39 14m27s 4m09s
ProofBench v1.1 $5.45 $0.48 13m43s 9m29s
Terminal-Bench Science $51.99 $6.09 1h59m 1h24m
Code Migration $112.10 $14.21 1h49m 1h24m
ProgramBench $75.68 $54.29 2h37m 5h28m
Terminal-Bench 4.0 $30.34 $6.65 1h08m 52m56s
Vibe Code Bench v1.1 $41.71 $2.54 1h02m 16m18s

Results available only for Claude Fable 5

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • MedCode
  • MedScribe
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SAGE
  • LiveCodeBench
  • SWE-bench
  • Public Benefits Bench v1.1

Results available only for Muse Spark 1.3 Max

  • MysteryMechanism
  • IOI
  • Vibe Code Bench 1-100
  • CyberBench v1.1
  • CUA-bench
Model details Claude Fable 5 Model details Muse Spark 1.3 Max