Claude Fable 5 vs Muse Spark 1.3: Benchmark Comparison
Claude Fable 5 has the higher score on 9 of 12 shared benchmarks; Muse Spark 1.3 leads on 3.
The largest observed score gap is 40.00 pts on ProofBench v1.1 , where Claude Fable 5 leads.
Reported ±1 standard-error ranges overlap on 1 of 12 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Claude Fable 5 | Muse Spark 1.3 | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 61.39% ±1.00 | 53.20% ±1.07 | 8.19 pts | Reported ±1 SE ranges do not overlap |
| Harvey's Legal Agent Benchmark | 11.25% ±2.15 | 22.92% ±3.42 | 11.67 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 49.52% ±3.48 | 40.87% ±3.42 | 8.65 pts | Reported ±1 SE ranges do not overlap |
| EMB | 73.67% ±2.47 | 62.71% ±2.96 | 10.95 pts | Reported ±1 SE ranges do not overlap |
| Finance Agent (v2) | 56.31% ±0.84 | 58.90% ±0.44 | 2.59 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 69.82% ±3.12 | 71.93% ±2.92 | 2.11 pts | Reported ±1 SE ranges overlap |
| ProofBench v1.1 | 95.00% ±2.19 | 55.00% ±5.00 | 40.00 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench Science | 15.71% ±4.38 | 4.29% ±2.44 | 11.43 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 55.06% ±4.61 | 27.58% ±4.07 | 27.48 pts | Reported ±1 SE ranges do not overlap |
| ProgramBench | 2.00% ±0.99 | 0.50% ±0.50 | 1.50 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench 4.0 | 41.41% ±1.34 | 10.61% ±1.75 | 30.81 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 90.35% ±2.10 | 82.86% ±2.90 | 7.49 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Claude Fable 5 average | Muse Spark 1.3 average |
|---|---|---|
| Legal | 30.38% | 31.89% |
| Finance | 66.60% | 64.52% |
| Math | 95.00% | 55.00% |
| Science | 15.71% | 4.29% |
| Coding | 47.21% | 30.39% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Claude Fable 5 cost | Muse Spark 1.3 cost | Claude Fable 5 latency | Muse Spark 1.3 latency |
|---|---|---|---|---|
| Vals Index | $29.59 | $3.38 | 42m24s | 32m42s |
| Harvey's Legal Agent Benchmark | $19.23 | $2.76 | 26m53s | 16m39s |
| Legal Research Bench | $9.79 | $0.65 | 22m18s | 8m51s |
| EMB | $12.35 | $3.55 | 26m34s | 22m55s |
| Finance Agent (v2) | $8.06 | $0.74 | 10m11s | 5m42s |
| Tax Agent Bench | $6.70 | $0.32 | 14m27s | 8m25s |
| ProofBench v1.1 | $5.45 | $0.43 | 13m43s | 8m54s |
| Terminal-Bench Science | $51.99 | $4.41 | 1h59m | 1h08m |
| Code Migration | $112.10 | $4.62 | 1h49m | 1h08m |
| ProgramBench | $75.68 | $11.56 | 2h37m | 1h09m |
| Terminal-Bench 4.0 | $30.34 | $7.04 | 1h08m | 1h59m |
| Vibe Code Bench v1.1 | $41.71 | $2.10 | 1h02m | 12m54s |
Results available only for Claude Fable 5
- Vals RSI Index
- LegalBench
- MortgageTax
- TaxEval v2
- MedCode
- MedScribe
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- SAGE
- LiveCodeBench
- SWE-bench
- Public Benefits Bench v1.1
Results available only for Muse Spark 1.3
- IOI
- CyberBench v1.1