Muse Spark 1.2 vs GPT 5.5: Benchmark Comparison

Muse Spark 1.2 has the higher score on 10 of 19 shared benchmarks; GPT 5.5 leads on 8.

The largest observed score gap is 21.67 pts on Harvey's Legal Agent Benchmark , where Muse Spark 1.2 leads.

Reported ±1 standard-error ranges overlap on 7 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Muse Spark 1.2 GPT 5.5 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 25.42% ±3.67 3.75% ±1.17 21.67 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 43.75% ±3.45 40.38% ±3.41 3.37 pts Reported ±1 SE ranges overlap
LegalBench 85.26% ±0.45 86.52% ±0.41 1.25 pts Reported ±1 SE ranges do not overlap
EMB 56.98% ±3.13 64.54% ±2.87 7.57 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 60.60% ±0.28 51.76% ±0.55 8.84 pts Reported ±1 SE ranges do not overlap
MortgageTax 65.42% ±0.93 68.76% ±0.91 3.34 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 56.86% ±2.32 60.46% ±3.22 3.60 pts Reported ±1 SE ranges overlap
TaxEval v2 80.38% ±0.76 74.98% ±0.85 5.40 pts Reported ±1 SE ranges do not overlap
MedCode 49.35% ±2.19 49.10% ±2.19 0.25 pts Reported ±1 SE ranges overlap
MedScribe 90.06% ±1.96 86.87% ±1.93 3.19 pts Reported ±1 SE ranges overlap
MMLU Pro 88.28% ±0.32 88.14% ±0.32 0.13 pts Reported ±1 SE ranges overlap
MMMU Pro 86.13% ±0.83 88.27% ±0.77 2.14 pts Reported ±1 SE ranges do not overlap
SAGE 47.66% ±3.44 51.53% ±3.95 3.87 pts Reported ±1 SE ranges overlap
Code Migration 29.95% ±4.02 45.16% ±4.16 15.21 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
SkillsBench 53.04% ±4.42 62.21% ±4.52 9.17 pts Reported ±1 SE ranges do not overlap
SWE-bench 86.60% ±1.52 82.60% ±1.70 4.00 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 79.10% ±3.31 69.85% ±4.54 9.25 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 68.47% ±1.21 60.89% ±1.27 7.58 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Muse Spark 1.2 average GPT 5.5 average
Legal 51.48% 43.55%
Finance 64.05% 64.10%
Healthcare 69.70% 67.98%
Academic 87.20% 88.21%
Education 47.66% 51.53%
Coding 49.84% 52.06%
Social Mobility 68.47% 60.89%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Muse Spark 1.2 cost GPT 5.5 cost Muse Spark 1.2 latency GPT 5.5 latency
Harvey's Legal Agent Benchmark $2.09 $4.60 24m16s 12m14s
Legal Research Bench $0.51 $7.40 5m07s 37m45s
LegalBench N/A N/A 30.10s 18.14s
EMB $2.31 $3.27 17m07s 15m11s
Finance Agent (v2) $0.77 $4.15 5m04s 11m02s
MortgageTax N/A N/A 44.62s 28.24s
Tax Agent Bench $0.27 $4.07 119.91s 15m52s
TaxEval v2 N/A N/A 48.31s 95.83s
MedCode N/A N/A 60.17s 2m40s
MedScribe N/A N/A 61.46s 2m13s
MMLU Pro N/A N/A 44.23s 42.11s
MMMU Pro N/A N/A 73.84s 54.15s
SAGE N/A N/A 29.07s 76.05s
Code Migration $3.78 $6.44 31m07s 31m35s
ProgramBench $2.47 $6.95 26m38s 22m03s
SkillsBench $0.65 $2.54 13m00s 6m55s
SWE-bench $0.55 $1.36 8m07s 7m06s
Vibe Code Bench v1.1 $1.53 $16.66 19m55s 31m52s
Public Benefits Bench v1.1 $0.59 $3.97 8m55s 39m37s

Results available only for Muse Spark 1.2

  • Vals Index
  • ProofBench v1.1
  • BioMysteryBench
  • IOI
  • Terminal-Bench 4.0
  • CyberBench v1.1

Results available only for GPT 5.5

  • Vals RSI Index
  • GPQA Diamond
  • LiveCodeBench
  • SRE Bench
  • Time Horizon Index: KSP
Model details Muse Spark 1.2 Model details GPT 5.5