Model comparison
Claude Sonnet 5.5 vs Muse Spark 1.3 Max: Benchmark Comparison
Claude Sonnet 5.5 has the higher score on 10 of 14 shared benchmarks; Muse Spark 1.3 Max leads on 4.
The largest observed score gap is 42.00 pts on ProofBench v1.1 , where Claude Sonnet 5.5 leads.
Reported ±1 standard-error ranges overlap on 2 of 14 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Claude Sonnet 5.5 | Muse Spark 1.3 Max | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 69.22% ±0.96 | 64.53% ±1.23 | 4.69 pts | Reported ±1 SE ranges do not overlap |
| Harvey's Legal Agent Benchmark | 2.92% ±1.36 | 23.75% ±3.71 | 20.83 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 48.08% ±3.47 | 55.29% ±3.46 | 7.21 pts | Reported ±1 SE ranges do not overlap |
| EMB | 75.71% ±2.44 | 67.43% ±3.06 | 8.29 pts | Reported ±1 SE ranges do not overlap |
| Finance Agent (v2) | 58.10% ±0.67 | 59.96% ±2.06 | 1.85 pts | Reported ±1 SE ranges overlap |
| Tax Agent Bench | 73.39% ±2.98 | 72.44% ±2.88 | 0.95 pts | Reported ±1 SE ranges overlap |
| ProofBench v1.1 | 100.00% ±0.00 | 58.00% ±4.96 | 42.00 pts | Reported ±1 SE ranges do not overlap |
| MysteryMechanism | 49.10% ±3.36 | 36.04% ±3.23 | 13.06 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench Science | 38.57% ±5.86 | 14.29% ±4.21 | 24.28 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 69.83% ±4.26 | 47.41% ±4.26 | 22.41 pts | Reported ±1 SE ranges do not overlap |
| IOI | 83.06% ±3.74 | 56.56% ±2.52 | 26.50 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench 4.0 | 53.03% ±1.51 | 27.78% ±2.20 | 25.25 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 92.39% ±1.26 | 85.86% ±2.51 | 6.54 pts | Reported ±1 SE ranges do not overlap |
| CyberBench v1.1 | 59.58% ±5.21 | 72.74% ±5.67 | 13.15 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Claude Sonnet 5.5 average | Muse Spark 1.3 Max average |
|---|---|---|
| Index | 69.22% | 64.53% |
| Legal | 25.50% | 39.52% |
| Finance | 69.07% | 66.61% |
| Math | 100.00% | 58.00% |
| Science | 43.83% | 25.16% |
| Coding | 74.58% | 54.40% |
| Beta | 59.58% | 72.74% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Claude Sonnet 5.5 cost | Muse Spark 1.3 Max cost | Claude Sonnet 5.5 latency | Muse Spark 1.3 Max latency |
|---|---|---|---|---|
| Vals Index | $20.80 | $3.40 | 1h10m | 20m30s |
| Harvey's Legal Agent Benchmark | $16.61 | $2.24 | 57m44s | 14m35s |
| Legal Research Bench | $14.83 | $0.59 | 1h33m | 5m17s |
| EMB | $8.79 | $2.55 | 36m45s | 13m50s |
| Finance Agent (v2) | $6.55 | $0.76 | 31m43s | 3m19s |
| Tax Agent Bench | $9.25 | $0.39 | 59m21s | 4m09s |
| ProofBench v1.1 | $0.56 | $0.48 | 5m48s | 9m29s |
| MysteryMechanism | $3.45 | $0.85 | 23m30s | 7m32s |
| Terminal-Bench Science | $31.98 | $9.87 | 3h22m | 2h39m |
| Code Migration | $75.83 | $14.21 | 3h20m | 1h24m |
| IOI | $7.52 | $1.73 | 45m51s | 28m58s |
| Terminal-Bench 4.0 | $19.33 | $5.38 | 2h06m | 3h03m |
| Vibe Code Bench v1.1 | $31.25 | $2.54 | 1h01m | 16m18s |
| CyberBench v1.1 | $4.42 | $3.34 | 33m53s | 16m18s |
Results available only for Claude Sonnet 5.5
- MedCode
- MedScribe
- BioMysteryBench
- SAGE
- SRE Bench
- Public Benefits Bench v1.1
Results available only for Muse Spark 1.3 Max
- Vals RSI Index
- Vibe Code Bench 1-100
- CUA-bench