Muse Spark 1.3 Max vs GPT-5.6 Sol: Benchmark Comparison

Muse Spark 1.3 Max has the higher score on 9 of 18 shared benchmarks; GPT-5.6 Sol leads on 9.

The largest observed score gap is 34.61 pts on IOI , where GPT-5.6 Sol leads.

Reported ±1 standard-error ranges overlap on 9 of 16 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Muse Spark 1.3 Max GPT-5.6 Sol Gap Reported uncertainty
Vals Index 58.16% ±1.19 58.01% ±1.02 0.16 pts Reported ±1 SE ranges overlap
Vals RSI Index 19.64% 23.88% 4.24 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 23.75% ±3.55 2.50% ±0.83 21.25 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 55.29% ±3.46 48.08% ±3.47 7.21 pts Reported ±1 SE ranges do not overlap
EMB 67.43% ±3.06 72.34% ±2.37 4.91 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 59.96% ±2.06 53.76% ±0.85 6.20 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 72.44% ±2.88 67.95% ±3.12 4.49 pts Reported ±1 SE ranges overlap
ProofBench v1.1 58.00% ±4.96 83.00% ±3.77 25.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 36.04% ±3.23 33.33% ±3.17 2.70 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 10.00% ±3.61 20.00% ±4.82 10.00 pts Reported ±1 SE ranges do not overlap
Code Migration 47.41% ±4.26 52.92% ±4.35 5.50 pts Reported ±1 SE ranges overlap
IOI 56.56% ±2.52 91.17% ±4.51 34.61 pts Reported ±1 SE ranges do not overlap
ProgramBench 2.50% ±1.11 1.50% ±0.86 1.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 24.75% ±0.51 37.88% ±0.00 13.13 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 20.46% ±3.85 20.01% ±3.63 0.45 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 85.86% ±2.51 80.50% ±3.72 5.36 pts Reported ±1 SE ranges overlap
CyberBench v1.1 72.74% ±5.67 76.31% ±5.18 3.57 pts Reported ±1 SE ranges overlap
CUA-bench 5.83% 8.33% 2.50 pts Uncertainty comparison unavailable

Performance by category

Category Muse Spark 1.3 Max average GPT-5.6 Sol average
Legal 39.52% 25.29%
Finance 66.61% 64.68%
Math 58.00% 83.00%
Science 23.02% 26.67%
Coding 39.59% 47.33%
Cyber 72.74% 76.31%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Muse Spark 1.3 Max cost GPT-5.6 Sol cost Muse Spark 1.3 Max latency GPT-5.6 Sol latency
Vals Index $3.79 $14.24 23m33s 37m23s
Vals RSI Index $48.82 $347.10 90h00m 90h00m
Harvey's Legal Agent Benchmark $2.26 $10.37 14m35s 22m50s
Legal Research Bench $0.59 $21.61 5m17s 1h17m
EMB $2.55 $6.01 13m50s 9m50s
Finance Agent (v2) $0.76 $1.25 3m19s 19m25s
Tax Agent Bench $0.39 $7.51 4m09s 44m41s
ProofBench v1.1 $0.48 $1.44 9m29s 8m40s
MysteryMechanism $0.85 $0.96 7m32s 8m43s
Terminal-Bench Science $6.09 $7.32 1h24m 1h44m
Code Migration $14.21 $24.54 1h24m 59m12s
IOI $1.73 $7.98 28m58s 1h02m
ProgramBench $54.29 $15.29 5h28m 29m24s
Terminal-Bench 4.0 $6.65 $7.98 52m56s 33m46s
Vibe Code Bench 1-100 $5.60 $47.43 20m02s 2h37m
Vibe Code Bench v1.1 $2.54 $33.40 16m18s 33m53s
CyberBench v1.1 $3.34 $3.05 16m18s 10m47s
CUA-bench $143.48 $206.99 36.39s 18.80s

Results available only for Muse Spark 1.3 Max

None.

Results available only for GPT-5.6 Sol

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • MedCode
  • MedScribe
  • BioMysteryBench
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SAGE
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • SRE Bench
  • Public Benefits Bench v1.1
  • Time Horizon Index: KSP
Model details Muse Spark 1.3 Max Model details GPT-5.6 Sol