Model comparison
Claude Sonnet 5 vs Muse Spark 1.3: Benchmark Comparison
Claude Sonnet 5 has the higher score on 5 of 13 shared benchmarks; Muse Spark 1.3 leads on 8.
The largest observed score gap is 22.00 pts on ProofBench v1.1 , where Claude Sonnet 5 leads.
Reported ±1 standard-error ranges overlap on 8 of 13 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Claude Sonnet 5 | Muse Spark 1.3 | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 51.77% ±1.09 | 53.16% ±1.07 | 1.39 pts | Reported ±1 SE ranges overlap |
| Harvey's Legal Agent Benchmark | 5.00% ±1.65 | 22.08% ±3.48 | 17.08 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 41.83% ±3.43 | 40.87% ±3.42 | 0.96 pts | Reported ±1 SE ranges overlap |
| EMB | 66.32% ±3.01 | 62.71% ±2.96 | 3.60 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 53.91% ±0.52 | 58.90% ±0.44 | 4.99 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 62.27% ±3.19 | 71.93% ±2.92 | 9.66 pts | Reported ±1 SE ranges do not overlap |
| ProofBench v1.1 | 77.00% ±4.23 | 55.00% ±5.00 | 22.00 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench Science | 2.86% ±2.01 | 4.29% ±2.44 | 1.43 pts | Reported ±1 SE ranges overlap |
| Code Migration | 44.39% ±4.25 | 27.58% ±4.07 | 16.81 pts | Reported ±1 SE ranges do not overlap |
| IOI | 45.00% ±2.75 | 43.94% ±1.45 | 1.06 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench 4.0 | 9.60% ±1.01 | 10.61% ±1.75 | 1.01 pts | Reported ±1 SE ranges overlap |
| Vibe Code Bench v1.1 | 81.33% ±3.05 | 82.86% ±2.90 | 1.53 pts | Reported ±1 SE ranges overlap |
| CyberBench v1.1 | 61.91% ±5.74 | 69.41% ±5.76 | 7.50 pts | Reported ±1 SE ranges overlap |
Performance by category
| Category | Claude Sonnet 5 average | Muse Spark 1.3 average |
|---|---|---|
| Legal | 23.41% | 31.47% |
| Finance | 60.83% | 64.52% |
| Math | 77.00% | 55.00% |
| Science | 2.86% | 4.29% |
| Coding | 45.08% | 41.25% |
| Beta | 61.91% | 69.41% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Claude Sonnet 5 cost | Muse Spark 1.3 cost | Claude Sonnet 5 latency | Muse Spark 1.3 latency |
|---|---|---|---|---|
| Vals Index | $13.72 | $3.37 | 54m16s | 32m42s |
| Harvey's Legal Agent Benchmark | $8.95 | $2.73 | 38m52s | 16m39s |
| Legal Research Bench | $2.72 | $0.65 | 25m46s | 8m51s |
| EMB | $10.29 | $3.55 | 49m46s | 22m55s |
| Finance Agent (v2) | $0.75 | $0.74 | 13m12s | 5m42s |
| Tax Agent Bench | $1.79 | $0.32 | 19m19s | 8m25s |
| ProofBench v1.1 | $1.37 | $0.43 | 14m58s | 8m54s |
| Terminal-Bench Science | $18.70 | $4.41 | 2h55m | 1h08m |
| Code Migration | $35.31 | $4.62 | 1h57m | 1h08m |
| IOI | $12.89 | $2.36 | 59m44s | 29m51s |
| Terminal-Bench 4.0 | $26.33 | $7.04 | 1h45m | 1h59m |
| Vibe Code Bench v1.1 | $25.39 | $2.10 | 1h07m | 12m54s |
| CyberBench v1.1 | $1.89 | $4.26 | 16m57s | 24m53s |
Results available only for Claude Sonnet 5
- LegalBench
- MortgageTax
- TaxEval v2
- MedCode
- MedScribe
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- SAGE
- LiveCodeBench
- ProgramBench
- SkillsBench
- SWE-bench
- Vibe Code Bench 1-100
- Public Benefits Bench v1.1
Results available only for Muse Spark 1.3
None.