Model comparison

Claude Sonnet 5.5 vs Muse Spark 1.3 Max: Benchmark Comparison

Claude Sonnet 5.5 has the higher score on 10 of 14 shared benchmarks; Muse Spark 1.3 Max leads on 4.

The largest observed score gap is 42.00 pts on ProofBench v1.1 , where Claude Sonnet 5.5 leads.

Reported ±1 standard-error ranges overlap on 2 of 14 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5.5 Muse Spark 1.3 Max Gap Reported uncertainty
Vals Index 69.22% ±0.96 64.53% ±1.23 4.69 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 2.92% ±1.36 23.75% ±3.71 20.83 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 48.08% ±3.47 55.29% ±3.46 7.21 pts Reported ±1 SE ranges do not overlap
EMB 75.71% ±2.44 67.43% ±3.06 8.29 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 58.10% ±0.67 59.96% ±2.06 1.85 pts Reported ±1 SE ranges overlap
Tax Agent Bench 73.39% ±2.98 72.44% ±2.88 0.95 pts Reported ±1 SE ranges overlap
ProofBench v1.1 100.00% ±0.00 58.00% ±4.96 42.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 49.10% ±3.36 36.04% ±3.23 13.06 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 38.57% ±5.86 14.29% ±4.21 24.28 pts Reported ±1 SE ranges do not overlap
Code Migration 69.83% ±4.26 47.41% ±4.26 22.41 pts Reported ±1 SE ranges do not overlap
IOI 83.06% ±3.74 56.56% ±2.52 26.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 53.03% ±1.51 27.78% ±2.20 25.25 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 92.39% ±1.26 85.86% ±2.51 6.54 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 59.58% ±5.21 72.74% ±5.67 13.15 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Sonnet 5.5 average Muse Spark 1.3 Max average
Index 69.22% 64.53%
Legal 25.50% 39.52%
Finance 69.07% 66.61%
Math 100.00% 58.00%
Science 43.83% 25.16%
Coding 74.58% 54.40%
Beta 59.58% 72.74%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5.5 cost Muse Spark 1.3 Max cost Claude Sonnet 5.5 latency Muse Spark 1.3 Max latency
Vals Index $20.80 $3.40 1h10m 20m30s
Harvey's Legal Agent Benchmark $16.61 $2.24 57m44s 14m35s
Legal Research Bench $14.83 $0.59 1h33m 5m17s
EMB $8.79 $2.55 36m45s 13m50s
Finance Agent (v2) $6.55 $0.76 31m43s 3m19s
Tax Agent Bench $9.25 $0.39 59m21s 4m09s
ProofBench v1.1 $0.56 $0.48 5m48s 9m29s
MysteryMechanism $3.45 $0.85 23m30s 7m32s
Terminal-Bench Science $31.98 $9.87 3h22m 2h39m
Code Migration $75.83 $14.21 3h20m 1h24m
IOI $7.52 $1.73 45m51s 28m58s
Terminal-Bench 4.0 $19.33 $5.38 2h06m 3h03m
Vibe Code Bench v1.1 $31.25 $2.54 1h01m 16m18s
CyberBench v1.1 $4.42 $3.34 33m53s 16m18s

Results available only for Claude Sonnet 5.5

  • MedCode
  • MedScribe
  • BioMysteryBench
  • SAGE
  • SRE Bench
  • Public Benefits Bench v1.1

Results available only for Muse Spark 1.3 Max

  • Vals RSI Index
  • Vibe Code Bench 1-100
  • CUA-bench
Model details Claude Sonnet 5.5 Model details Muse Spark 1.3 Max