Muse Spark 1.3 vs GPT-5.6 Luna: Benchmark Comparison

Muse Spark 1.3 has the higher score on 8 of 14 shared benchmarks; GPT-5.6 Luna leads on 6.

The largest observed score gap is 21.67 pts on Harvey's Legal Agent Benchmark , where Muse Spark 1.3 leads.

Reported ±1 standard-error ranges overlap on 8 of 14 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Muse Spark 1.3 GPT-5.6 Luna Gap Reported uncertainty
Vals Index 53.20% ±1.07 51.69% ±1.06 1.51 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 22.92% ±3.42 1.25% ±0.83 21.67 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 40.87% ±3.42 36.54% ±3.35 4.33 pts Reported ±1 SE ranges overlap
EMB 62.71% ±2.96 67.12% ±2.94 4.40 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 58.90% ±0.44 55.04% ±0.31 3.86 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 71.93% ±2.92 60.81% ±3.26 11.12 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 55.00% ±5.00 60.00% ±4.92 5.00 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 4.29% ±2.44 0.00% ±0.00 4.29 pts Reported ±1 SE ranges do not overlap
Code Migration 27.58% ±4.07 44.55% ±4.24 16.97 pts Reported ±1 SE ranges do not overlap
IOI 43.94% ±1.45 61.78% ±11.61 17.83 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.00% ±0.00 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 10.61% ±1.75 11.62% ±1.01 1.01 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 82.86% ±2.90 77.06% ±3.07 5.80 pts Reported ±1 SE ranges overlap
CyberBench v1.1 69.41% ±5.76 73.63% ±5.56 4.23 pts Reported ±1 SE ranges overlap

Performance by category

Category Muse Spark 1.3 average GPT-5.6 Luna average
Legal 31.89% 18.89%
Finance 64.52% 60.99%
Math 55.00% 60.00%
Science 4.29% 0.00%
Coding 33.10% 39.00%
Cyber 69.41% 73.63%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Muse Spark 1.3 cost GPT-5.6 Luna cost Muse Spark 1.3 latency GPT-5.6 Luna latency
Vals Index $3.38 $0.82 32m42s 29m27s
Harvey's Legal Agent Benchmark $2.76 $0.38 16m39s 10m56s
Legal Research Bench $0.65 $0.85 8m51s 39m34s
EMB $3.55 $0.37 22m55s 11m00s
Finance Agent (v2) $0.74 $0.28 5m42s 12m52s
Tax Agent Bench $0.32 $0.47 8m25s 42m46s
ProofBench v1.1 $0.43 $0.12 8m54s 8m34s
Terminal-Bench Science $4.41 $0.58 1h08m 2h03m
Code Migration $4.62 $1.88 1h08m 58m35s
IOI $2.36 $0.58 29m51s 1h10m
ProgramBench $11.56 $0.94 1h09m 41m34s
Terminal-Bench 4.0 $7.04 $0.73 1h59m 35m55s
Vibe Code Bench v1.1 $2.10 $0.73 12m54s 25m52s
CyberBench v1.1 $4.26 $0.40 24m53s 16m15s

Results available only for Muse Spark 1.3

None.

Results available only for GPT-5.6 Luna

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • MedCode
  • MedScribe
  • BioMysteryBench
  • MysteryMechanism
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SAGE
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
  • Public Benefits Bench v1.1
Model details Muse Spark 1.3 Model details GPT-5.6 Luna