Model comparison

Claude Sonnet 5 vs Muse Spark 1.3: Benchmark Comparison

Claude Sonnet 5 has the higher score on 5 of 13 shared benchmarks; Muse Spark 1.3 leads on 8.

The largest observed score gap is 22.00 pts on ProofBench v1.1 , where Claude Sonnet 5 leads.

Reported ±1 standard-error ranges overlap on 8 of 13 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5 Muse Spark 1.3 Gap Reported uncertainty
Vals Index 51.77% ±1.09 53.16% ±1.07 1.39 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 5.00% ±1.65 22.08% ±3.48 17.08 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.83% ±3.43 40.87% ±3.42 0.96 pts Reported ±1 SE ranges overlap
EMB 66.32% ±3.01 62.71% ±2.96 3.60 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.91% ±0.52 58.90% ±0.44 4.99 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.27% ±3.19 71.93% ±2.92 9.66 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 77.00% ±4.23 55.00% ±5.00 22.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 2.86% ±2.01 4.29% ±2.44 1.43 pts Reported ±1 SE ranges overlap
Code Migration 44.39% ±4.25 27.58% ±4.07 16.81 pts Reported ±1 SE ranges do not overlap
IOI 45.00% ±2.75 43.94% ±1.45 1.06 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 9.60% ±1.01 10.61% ±1.75 1.01 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 81.33% ±3.05 82.86% ±2.90 1.53 pts Reported ±1 SE ranges overlap
CyberBench v1.1 61.91% ±5.74 69.41% ±5.76 7.50 pts Reported ±1 SE ranges overlap

Performance by category

Category Claude Sonnet 5 average Muse Spark 1.3 average
Legal 23.41% 31.47%
Finance 60.83% 64.52%
Math 77.00% 55.00%
Science 2.86% 4.29%
Coding 45.08% 41.25%
Beta 61.91% 69.41%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5 cost Muse Spark 1.3 cost Claude Sonnet 5 latency Muse Spark 1.3 latency
Vals Index $13.72 $3.37 54m16s 32m42s
Harvey's Legal Agent Benchmark $8.95 $2.73 38m52s 16m39s
Legal Research Bench $2.72 $0.65 25m46s 8m51s
EMB $10.29 $3.55 49m46s 22m55s
Finance Agent (v2) $0.75 $0.74 13m12s 5m42s
Tax Agent Bench $1.79 $0.32 19m19s 8m25s
ProofBench v1.1 $1.37 $0.43 14m58s 8m54s
Terminal-Bench Science $18.70 $4.41 2h55m 1h08m
Code Migration $35.31 $4.62 1h57m 1h08m
IOI $12.89 $2.36 59m44s 29m51s
Terminal-Bench 4.0 $26.33 $7.04 1h45m 1h59m
Vibe Code Bench v1.1 $25.39 $2.10 1h07m 12m54s
CyberBench v1.1 $1.89 $4.26 16m57s 24m53s

Results available only for Claude Sonnet 5

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • MedCode
  • MedScribe
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SAGE
  • LiveCodeBench
  • ProgramBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
  • Public Benefits Bench v1.1

Results available only for Muse Spark 1.3

None.

Model details Claude Sonnet 5 Model details Muse Spark 1.3