Claude Opus 5 vs Muse Spark 1.3 Max: Benchmark Comparison

Claude Opus 5 has the higher score on 14 of 18 shared benchmarks; Muse Spark 1.3 Max leads on 3.

The largest observed score gap is 41.00 pts on ProofBench v1.1 , where Claude Opus 5 leads.

Reported ±1 standard-error ranges overlap on 8 of 16 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Opus 5 Muse Spark 1.3 Max Gap Reported uncertainty
Vals Index 63.67% ±0.96 58.16% ±1.19 5.51 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 33.02% 19.64% 13.38 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 6.67% ±1.65 23.75% ±3.55 17.08 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 55.29% ±3.46 55.29% ±3.46 0.00 pts Reported ±1 SE ranges overlap
EMB 73.56% ±2.24 67.43% ±3.06 6.13 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 58.63% ±0.08 59.96% ±2.06 1.32 pts Reported ±1 SE ranges overlap
Tax Agent Bench 75.06% ±2.96 72.44% ±2.88 2.62 pts Reported ±1 SE ranges overlap
ProofBench v1.1 99.00% ±1.00 58.00% ±4.96 41.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 37.39% ±3.25 36.04% ±3.23 1.35 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 27.14% ±5.35 10.00% ±3.61 17.14 pts Reported ±1 SE ranges do not overlap
Code Migration 57.47% ±4.37 47.41% ±4.26 10.06 pts Reported ±1 SE ranges do not overlap
IOI 84.33% ±9.96 56.56% ±2.52 27.78 pts Reported ±1 SE ranges do not overlap
ProgramBench 3.00% ±1.21 2.50% ±1.11 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 53.53% ±1.34 24.75% ±0.51 28.79 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 28.53% ±4.33 20.46% ±3.85 8.07 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 88.40% ±3.00 85.86% ±2.51 2.55 pts Reported ±1 SE ranges overlap
CyberBench v1.1 65.36% ±5.55 72.74% ±5.67 7.38 pts Reported ±1 SE ranges overlap
CUA-bench 9.00% 5.83% 3.17 pts Uncertainty comparison unavailable

Performance by category

Category Claude Opus 5 average Muse Spark 1.3 Max average
Legal 30.98% 39.52%
Finance 69.08% 66.61%
Math 99.00% 58.00%
Science 32.27% 23.02%
Coding 52.55% 39.59%
Cyber 65.36% 72.74%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Opus 5 cost Muse Spark 1.3 Max cost Claude Opus 5 latency Muse Spark 1.3 Max latency
Vals Index $19.31 $3.79 57m41s 23m33s
Vals RSI Index $418.69 $48.82 90h00m 90h00m
Harvey's Legal Agent Benchmark $23.67 $2.26 55m37s 14m35s
Legal Research Bench $6.76 $0.59 31m01s 5m17s
EMB $6.07 $2.55 25m59s 13m50s
Finance Agent (v2) $5.12 $0.76 9m58s 3m19s
Tax Agent Bench $5.14 $0.39 16m48s 4m09s
ProofBench v1.1 $2.11 $0.48 9m51s 9m29s
MysteryMechanism $3.28 $0.85 16m53s 7m32s
Terminal-Bench Science $32.54 $6.09 2h12m 1h24m
Code Migration $60.51 $14.21 2h51m 1h24m
IOI $16.48 $1.73 56m42s 28m58s
ProgramBench $60.29 $54.29 2h55m 5h28m
Terminal-Bench 4.0 $18.60 $6.65 1h07m 52m56s
Vibe Code Bench 1-100 $41.49 $5.60 1h22m 20m02s
Vibe Code Bench v1.1 $33.88 $2.54 1h28m 16m18s
CyberBench v1.1 $2.87 $3.34 19m25s 16m18s
CUA-bench $232.23 $143.48 25.29s 36.39s

Results available only for Claude Opus 5

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • MedCode
  • MedScribe
  • BioMysteryBench
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SAGE
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • SRE Bench
  • Public Benefits Bench v1.1
  • Time Horizon Index: KSP

Results available only for Muse Spark 1.3 Max

None.

Model details Claude Opus 5 Model details Muse Spark 1.3 Max