Claude Opus 5.5 vs Muse Spark 1.3 Max: Benchmark Comparison

Claude Opus 5.5 has the higher score on 13 of 18 shared benchmarks; Muse Spark 1.3 Max leads on 5.

The largest observed score gap is 42.00 pts on ProofBench v1.1 , where Claude Opus 5.5 leads.

Reported ±1 standard-error ranges overlap on 3 of 16 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Opus 5.5 Muse Spark 1.3 Max Gap Reported uncertainty
Vals Index 66.97% ±0.89 58.16% ±1.19 8.81 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 37.31% 19.64% 17.67 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 3.75% ±1.47 23.75% ±3.55 20.00 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 50.48% ±3.48 55.29% ±3.46 4.81 pts Reported ±1 SE ranges overlap
EMB 75.94% ±2.38 67.43% ±3.06 8.51 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 58.59% ±0.17 59.96% ±2.06 1.37 pts Reported ±1 SE ranges overlap
Tax Agent Bench 70.50% ±3.15 72.44% ±2.88 1.94 pts Reported ±1 SE ranges overlap
ProofBench v1.1 100.00% ±0.00 58.00% ±4.96 42.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 49.55% ±3.36 36.04% ±3.23 13.51 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 47.14% ±6.01 10.00% ±3.61 37.14 pts Reported ±1 SE ranges do not overlap
Code Migration 66.65% ±4.33 47.41% ±4.26 19.23 pts Reported ±1 SE ranges do not overlap
IOI 95.06% ±4.94 56.56% ±2.52 38.50 pts Reported ±1 SE ranges do not overlap
ProgramBench 18.50% ±2.75 2.50% ±1.11 16.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 65.15% ±0.00 24.75% ±0.51 40.41 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 30.36% ±4.83 20.46% ±3.85 9.90 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 90.29% ±1.53 85.86% ±2.51 4.44 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 55.36% ±5.13 72.74% ±5.67 17.38 pts Reported ±1 SE ranges do not overlap
CUA-bench 14.00% 5.83% 8.17 pts Uncertainty comparison unavailable

Performance by category

Category Claude Opus 5.5 average Muse Spark 1.3 Max average
Legal 27.12% 39.52%
Finance 68.34% 66.61%
Math 100.00% 58.00%
Science 48.35% 23.02%
Coding 61.00% 39.59%
Cyber 55.36% 72.74%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Opus 5.5 cost Muse Spark 1.3 Max cost Claude Opus 5.5 latency Muse Spark 1.3 Max latency
Vals Index $32.14 $3.79 1h19m 23m33s
Vals RSI Index $411.86 $48.82 90h00m 90h00m
Harvey's Legal Agent Benchmark $21.38 $2.26 49m00s 14m35s
Legal Research Bench $25.16 $0.59 1h52m 5m17s
EMB $11.02 $2.55 35m06s 13m50s
Finance Agent (v2) $9.22 $0.76 33m08s 3m19s
Tax Agent Bench $15.09 $0.39 1h07m 4m09s
ProofBench v1.1 $0.96 $0.48 5m51s 9m29s
MysteryMechanism $4.62 $0.85 19m38s 7m32s
Terminal-Bench Science $19.12 $6.09 2h28m 1h24m
Code Migration $112.97 $14.21 3h19m 1h24m
IOI $5.26 $1.73 17m15s 28m58s
ProgramBench $68.05 $54.29 2h24m 5h28m
Terminal-Bench 4.0 $13.20 $6.65 1h04m 52m56s
Vibe Code Bench 1-100 $44.12 $5.60 2h28m 20m02s
Vibe Code Bench v1.1 $57.92 $2.54 1h36m 16m18s
CyberBench v1.1 $3.48 $3.34 18m11s 16m18s
CUA-bench $188.04 $143.48 14.71s 36.39s

Results available only for Claude Opus 5.5

  • MedCode
  • MedScribe
  • BioMysteryBench
  • SAGE
  • SRE Bench
  • Public Benefits Bench v1.1
  • Time Horizon Index: KSP

Results available only for Muse Spark 1.3 Max

None.

Model details Claude Opus 5.5 Model details Muse Spark 1.3 Max