Muse Spark 1.2 vs GPT-6 Luna: Benchmark Comparison

Muse Spark 1.2 has the higher score on 7 of 18 shared benchmarks; GPT-6 Luna leads on 10.

The largest observed score gap is 33.78 pts on IOI , where GPT-6 Luna leads.

Reported ±1 standard-error ranges overlap on 7 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Muse Spark 1.2 GPT-6 Luna Gap Reported uncertainty
Vals Index 49.29% ±1.10 51.22% ±1.07 1.93 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 25.42% ±3.67 2.92% ±1.36 22.50 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 43.75% ±3.45 30.29% ±3.19 13.46 pts Reported ±1 SE ranges do not overlap
EMB 56.98% ±3.13 68.52% ±2.81 11.54 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 60.60% ±0.28 49.87% ±0.23 10.73 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 56.86% ±2.32 58.86% ±3.21 2.00 pts Reported ±1 SE ranges overlap
MedCode 49.35% ±2.19 44.69% ±2.30 4.66 pts Reported ±1 SE ranges do not overlap
MedScribe 90.06% ±1.96 83.71% ±1.95 6.35 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 43.00% ±4.98 64.00% ±4.82 21.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 64.81% ±1.33 61.48% ±2.59 3.33 pts Reported ±1 SE ranges overlap
SAGE 47.66% ±3.44 48.09% ±3.38 0.43 pts Reported ±1 SE ranges overlap
Code Migration 29.95% ±4.02 42.55% ±4.42 12.60 pts Reported ±1 SE ranges do not overlap
IOI 21.78% ±0.87 55.56% ±8.87 33.78 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 6.06% ±1.51 13.64% ±1.51 7.57 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 79.10% ±3.31 81.65% ±3.38 2.55 pts Reported ±1 SE ranges overlap
CyberBench v1.1 69.52% ±5.56 76.25% ±5.29 6.73 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 68.47% ±1.21 57.65% ±1.28 10.83 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Muse Spark 1.2 average GPT-6 Luna average
Legal 34.58% 16.60%
Finance 58.15% 59.08%
Healthcare 69.70% 64.20%
Math 43.00% 64.00%
Science 64.81% 61.48%
Education 47.66% 48.09%
Coding 27.48% 38.78%
Cyber 69.52% 76.25%
Social Mobility 68.47% 57.65%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Muse Spark 1.2 cost GPT-6 Luna cost Muse Spark 1.2 latency GPT-6 Luna latency
Vals Index $1.94 $0.43 22m56s 30m34s
Harvey's Legal Agent Benchmark $2.09 $0.30 24m16s 17m32s
Legal Research Bench $0.51 $0.44 5m07s 42m28s
EMB $2.31 $0.14 17m07s 17m21s
Finance Agent (v2) $0.77 $0.12 5m04s 14m58s
Tax Agent Bench $0.27 $0.19 119.91s 33m16s
MedCode N/A N/A 60.17s 108.77s
MedScribe N/A N/A 61.46s 2m50s
ProofBench v1.1 $0.44 $0.04 7m49s 7m31s
BioMysteryBench $0.99 $0.05 23m37s 8m44s
SAGE N/A N/A 29.07s 91.07s
Code Migration $3.78 $0.60 31m07s 50m23s
IOI $2.64 $0.20 25m35s 35m47s
ProgramBench $2.47 $0.18 26m38s 28m55s
Terminal-Bench 4.0 $4.14 $0.35 1h18m 32m46s
Vibe Code Bench v1.1 $1.53 $1.35 19m55s 37m00s
CyberBench v1.1 $2.15 $0.13 20m21s 15m50s
Public Benefits Bench v1.1 $0.59 $0.22 8m55s 52m59s

Results available only for Muse Spark 1.2

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • MMLU Pro
  • MMMU Pro
  • SkillsBench
  • SWE-bench

Results available only for GPT-6 Luna

  • MysteryMechanism
  • Terminal-Bench Science
Model details Muse Spark 1.2 Model details GPT-6 Luna