Claude Opus 4.8 vs Muse Spark 1.3: Benchmark Comparison

Claude Opus 4.8 has the higher score on 6 of 11 shared benchmarks; Muse Spark 1.3 leads on 4.

The largest observed score gap is 19.67 pts on Code Migration , where Claude Opus 4.8 leads.

Reported ±1 standard-error ranges overlap on 5 of 11 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Opus 4.8 Muse Spark 1.3 Gap Reported uncertainty
Vals Index 55.10% ±1.01 53.20% ±1.07 1.90 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 9.58% ±2.15 22.92% ±3.42 13.33 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 43.75% ±3.45 40.87% ±3.42 2.88 pts Reported ±1 SE ranges overlap
EMB 69.37% ±2.61 62.71% ±2.96 6.66 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 53.92% ±0.16 58.90% ±0.44 4.98 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 64.91% ±3.17 71.93% ±2.92 7.03 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 4.29% ±2.44 4.29% ±2.44 0.00 pts Reported ±1 SE ranges overlap
Code Migration 47.25% ±4.18 27.58% ±4.07 19.67 pts Reported ±1 SE ranges do not overlap
ProgramBench 1.00% ±0.70 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 23.23% ±1.34 10.61% ±1.75 12.63 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 82.72% ±3.08 82.86% ±2.90 0.13 pts Reported ±1 SE ranges overlap

Performance by category

Category Claude Opus 4.8 average Muse Spark 1.3 average
Legal 26.67% 31.89%
Finance 62.73% 64.52%
Science 4.29% 4.29%
Coding 38.55% 30.39%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Opus 4.8 cost Muse Spark 1.3 cost Claude Opus 4.8 latency Muse Spark 1.3 latency
Vals Index $13.14 $3.38 39m48s 32m42s
Harvey's Legal Agent Benchmark $10.22 $2.76 24m08s 16m39s
Legal Research Bench $2.82 $0.65 11m54s 8m51s
EMB $12.06 $3.55 43m52s 22m55s
Finance Agent (v2) $4.22 $0.74 9m07s 5m42s
Tax Agent Bench $1.69 $0.32 7m59s 8m25s
Terminal-Bench Science $23.14 $4.41 1h55m 1h08m
Code Migration $30.51 $4.62 1h15m 1h08m
ProgramBench $31.27 $11.56 1h32m 1h09m
Terminal-Bench 4.0 $17.14 $7.04 1h11m 1h59m
Vibe Code Bench v1.1 $26.88 $2.10 1h16m 12m54s

Results available only for Claude Opus 4.8

  • Vals RSI Index
  • LegalBench
  • MortgageTax
  • TaxEval v2
  • MedCode
  • MedScribe
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SAGE
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Public Benefits Bench v1.1
  • Time Horizon Index: KSP

Results available only for Muse Spark 1.3

  • ProofBench v1.1
  • IOI
  • CyberBench v1.1
Model details Claude Opus 4.8 Model details Muse Spark 1.3