Muse Spark 1.3 Max vs GPT-6 Sol: Benchmark Comparison

Muse Spark 1.3 Max has the higher score on 7 of 16 shared benchmarks; GPT-6 Sol leads on 9.

The largest observed score gap is 26.44 pts on Legal Research Bench , where Muse Spark 1.3 Max leads.

Reported ±1 standard-error ranges overlap on 6 of 15 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Muse Spark 1.3 Max GPT-6 Sol Gap Reported uncertainty
Vals Index 58.16% ±1.19 57.54% ±1.01 0.63 pts Reported ±1 SE ranges overlap
Vals RSI Index 19.64% 28.11% 8.47 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 23.75% ±3.55 1.67% ±0.82 22.08 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 55.29% ±3.46 28.85% ±3.15 26.44 pts Reported ±1 SE ranges do not overlap
EMB 67.43% ±3.06 71.53% ±2.35 4.10 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 59.96% ±2.06 49.05% ±0.58 10.91 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 72.44% ±2.88 53.05% ±3.29 19.40 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 58.00% ±4.96 83.00% ±3.77 25.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 36.04% ±3.23 30.18% ±3.09 5.86 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 10.00% ±3.61 30.00% ±5.52 20.00 pts Reported ±1 SE ranges do not overlap
Code Migration 47.41% ±4.26 57.20% ±4.21 9.78 pts Reported ±1 SE ranges do not overlap
IOI 56.56% ±2.52 82.61% ±9.24 26.06 pts Reported ±1 SE ranges do not overlap
ProgramBench 2.50% ±1.11 2.00% ±0.99 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 24.75% ±0.51 44.44% ±3.54 19.70 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 85.86% ±2.51 87.82% ±2.53 1.97 pts Reported ±1 SE ranges overlap
CyberBench v1.1 72.74% ±5.67 77.98% ±5.11 5.24 pts Reported ±1 SE ranges overlap

Performance by category

Category Muse Spark 1.3 Max average GPT-6 Sol average
Legal 39.52% 15.26%
Finance 66.61% 57.88%
Math 58.00% 83.00%
Science 23.02% 30.09%
Coding 43.41% 54.81%
Cyber 72.74% 77.98%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Muse Spark 1.3 Max cost GPT-6 Sol cost Muse Spark 1.3 Max latency GPT-6 Sol latency
Vals Index $3.79 $7.58 23m33s 29m11s
Vals RSI Index $48.82 $665.11 90h00m 90h00m
Harvey's Legal Agent Benchmark $2.26 $3.36 14m35s 13m40s
Legal Research Bench $0.59 $4.90 5m17s 24m09s
EMB $2.55 $1.28 13m50s 10m22s
Finance Agent (v2) $0.76 $2.12 3m19s 9m00s
Tax Agent Bench $0.39 $2.28 4m09s 17m32s
ProofBench v1.1 $0.48 $0.48 9m29s 5m05s
MysteryMechanism $0.85 $0.46 7m32s 6m13s
Terminal-Bench Science $6.09 $5.82 1h24m 1h22m
Code Migration $14.21 $15.68 1h24m 1h26m
IOI $1.73 $2.70 28m58s 30m38s
ProgramBench $54.29 $11.67 5h28m 46m11s
Terminal-Bench 4.0 $6.65 $5.79 52m56s 34m56s
Vibe Code Bench v1.1 $2.54 $26.36 16m18s 38m56s
CyberBench v1.1 $3.34 $1.34 16m18s 12m14s

Results available only for Muse Spark 1.3 Max

  • Vibe Code Bench 1-100
  • CUA-bench

Results available only for GPT-6 Sol

  • MedCode
  • MedScribe
  • BioMysteryBench
  • SAGE
  • Public Benefits Bench v1.1
Model details Muse Spark 1.3 Max Model details GPT-6 Sol