Muse Spark 1.2 vs GLM 5.3: Benchmark Comparison

Muse Spark 1.2 has the higher score on 10 of 21 shared benchmarks; GLM 5.3 leads on 11.

The largest observed score gap is 46.67 pts on IOI , where GLM 5.3 leads.

Reported ±1 standard-error ranges overlap on 10 of 21 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Muse Spark 1.2 GLM 5.3 Gap Reported uncertainty
Vals Index 49.29% ±1.10 53.51% ±1.30 4.23 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 25.42% ±3.67 8.33% ±2.00 17.08 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 43.75% ±3.45 49.04% ±3.48 5.29 pts Reported ±1 SE ranges overlap
LegalBench 85.26% ±0.45 84.84% ±0.40 0.42 pts Reported ±1 SE ranges overlap
EMB 56.98% ±3.13 56.34% ±3.32 0.63 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 60.60% ±0.28 55.84% ±2.07 4.76 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 56.86% ±2.32 73.09% ±2.97 16.23 pts Reported ±1 SE ranges do not overlap
TaxEval v2 80.38% ±0.76 72.36% ±0.88 8.01 pts Reported ±1 SE ranges do not overlap
MedCode 49.35% ±2.19 42.86% ±2.11 6.48 pts Reported ±1 SE ranges do not overlap
MedScribe 90.06% ±1.96 88.81% ±2.00 1.25 pts Reported ±1 SE ranges overlap
ProofBench v1.1 43.00% ±4.98 49.00% ±5.02 6.00 pts Reported ±1 SE ranges overlap
MMLU Pro 88.28% ±0.32 86.77% ±0.34 1.51 pts Reported ±1 SE ranges do not overlap
Code Migration 29.95% ±4.02 44.22% ±4.29 14.27 pts Reported ±1 SE ranges do not overlap
IOI 21.78% ±0.87 68.44% ±7.48 46.67 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 1.50% ±0.86 1.00 pts Reported ±1 SE ranges overlap
SkillsBench 53.04% ±4.42 47.51% ±4.52 5.53 pts Reported ±1 SE ranges overlap
SWE-bench 86.60% ±1.52 95.40% ±0.94 8.80 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 6.06% ±1.51 38.89% ±1.82 32.83 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 79.10% ±3.31 78.13% ±3.94 0.97 pts Reported ±1 SE ranges overlap
CyberBench v1.1 69.52% ±5.56 72.08% ±5.41 2.56 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 68.47% ±1.21 68.54% ±1.21 0.07 pts Reported ±1 SE ranges overlap

Performance by category

Category Muse Spark 1.2 average GLM 5.3 average
Legal 51.48% 47.40%
Finance 63.70% 64.41%
Healthcare 69.70% 65.84%
Math 43.00% 49.00%
Academic 88.28% 86.77%
Coding 39.58% 53.44%
Cyber 69.52% 72.08%
Social Mobility 68.47% 68.54%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Muse Spark 1.2 cost GLM 5.3 cost Muse Spark 1.2 latency GLM 5.3 latency
Vals Index $1.94 $7.25 22m56s 1h13m
Harvey's Legal Agent Benchmark $2.09 $4.30 24m16s 42m38s
Legal Research Bench $0.51 $2.24 5m07s 51m50s
LegalBench N/A N/A 30.10s 22.48s
EMB $2.31 $3.79 17m07s 36m12s
Finance Agent (v2) $0.77 $1.07 5m04s 15m50s
Tax Agent Bench $0.27 $1.80 119.91s 50m53s
TaxEval v2 N/A N/A 48.31s 98.32s
MedCode N/A N/A 60.17s 2m49s
MedScribe N/A N/A 61.46s 2m02s
ProofBench v1.1 $0.44 $2.08 7m49s 42m23s
MMLU Pro N/A N/A 44.23s 61.78s
Code Migration $3.78 $24.91 31m07s 3h47m
IOI $2.64 $7.67 25m35s 1h46m
ProgramBench $2.47 $21.96 26m38s 4h16m
SkillsBench $0.65 $0.71 13m00s 13m07s
SWE-bench $0.55 $0.34 8m07s 14m01s
Terminal-Bench 4.0 $4.14 $9.37 1h18m 1h23m
Vibe Code Bench v1.1 $1.53 $12.45 19m55s 1h04m
CyberBench v1.1 $2.15 $2.69 20m21s 30m18s
Public Benefits Bench v1.1 $0.59 $0.95 8m55s 35m51s

Results available only for Muse Spark 1.2

  • MortgageTax
  • BioMysteryBench
  • MMMU Pro
  • SAGE

Results available only for GLM 5.3

  • Vals RSI Index
  • MysteryMechanism
  • Terminal-Bench Science
  • GPQA Diamond
  • LiveCodeBench
  • Vibe Code Bench 1-100
Model details Muse Spark 1.2 Model details GLM 5.3