Claude Opus 5 vs Muse Spark 1.3: Benchmark Comparison

Claude Opus 5 has the higher score on 11 of 14 shared benchmarks; Muse Spark 1.3 leads on 3.

The largest observed score gap is 44.00 pts on ProofBench v1.1 , where Claude Opus 5 leads.

Reported ±1 standard-error ranges overlap on 4 of 14 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Opus 5 Muse Spark 1.3 Gap Reported uncertainty
Vals Index 63.67% ±0.96 53.20% ±1.07 10.48 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 6.67% ±1.65 22.92% ±3.42 16.25 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 55.29% ±3.46 40.87% ±3.42 14.42 pts Reported ±1 SE ranges do not overlap
EMB 73.56% ±2.24 62.71% ±2.96 10.84 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 58.63% ±0.08 58.90% ±0.44 0.27 pts Reported ±1 SE ranges overlap
Tax Agent Bench 75.06% ±2.96 71.93% ±2.92 3.13 pts Reported ±1 SE ranges overlap
ProofBench v1.1 99.00% ±1.00 55.00% ±5.00 44.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 27.14% ±5.35 4.29% ±2.44 22.86 pts Reported ±1 SE ranges do not overlap
Code Migration 57.47% ±4.37 27.58% ±4.07 29.89 pts Reported ±1 SE ranges do not overlap
IOI 84.33% ±9.96 43.94% ±1.45 40.39 pts Reported ±1 SE ranges do not overlap
ProgramBench 3.00% ±1.21 0.50% ±0.50 2.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 53.53% ±1.34 10.61% ±1.75 42.93 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 88.40% ±3.00 82.86% ±2.90 5.55 pts Reported ±1 SE ranges overlap
CyberBench v1.1 65.36% ±5.55 69.41% ±5.76 4.05 pts Reported ±1 SE ranges overlap

Performance by category

Category Claude Opus 5 average Muse Spark 1.3 average
Legal 30.98% 31.89%
Finance 69.08% 64.52%
Math 99.00% 55.00%
Science 27.14% 4.29%
Coding 57.35% 33.10%
Cyber 65.36% 69.41%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Opus 5 cost Muse Spark 1.3 cost Claude Opus 5 latency Muse Spark 1.3 latency
Vals Index $19.31 $3.38 57m41s 32m42s
Harvey's Legal Agent Benchmark $23.67 $2.76 55m37s 16m39s
Legal Research Bench $6.76 $0.65 31m01s 8m51s
EMB $6.07 $3.55 25m59s 22m55s
Finance Agent (v2) $5.12 $0.74 9m58s 5m42s
Tax Agent Bench $5.14 $0.32 16m48s 8m25s
ProofBench v1.1 $2.11 $0.43 9m51s 8m54s
Terminal-Bench Science $32.54 $4.41 2h12m 1h08m
Code Migration $60.51 $4.62 2h51m 1h08m
IOI $16.48 $2.36 56m42s 29m51s
ProgramBench $60.29 $11.56 2h55m 1h09m
Terminal-Bench 4.0 $18.60 $7.04 1h07m 1h59m
Vibe Code Bench v1.1 $33.88 $2.10 1h28m 12m54s
CyberBench v1.1 $2.87 $4.26 19m25s 24m53s

Results available only for Claude Opus 5

  • Vals RSI Index
  • LegalBench
  • MortgageTax
  • TaxEval v2
  • MedCode
  • MedScribe
  • BioMysteryBench
  • MysteryMechanism
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SAGE
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
  • SRE Bench
  • Public Benefits Bench v1.1
  • CUA-bench
  • Time Horizon Index: KSP

Results available only for Muse Spark 1.3

None.

Model details Claude Opus 5 Model details Muse Spark 1.3