Muse Spark 1.2 vs GLM 5.3: Benchmark Comparison
Muse Spark 1.2 has the higher score on 10 of 21 shared benchmarks; GLM 5.3 leads on 11.
The largest observed score gap is 46.67 pts on IOI , where GLM 5.3 leads.
Reported ±1 standard-error ranges overlap on 10 of 21 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Muse Spark 1.2 | GLM 5.3 | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 49.29% ±1.10 | 53.51% ±1.30 | 4.23 pts | Reported ±1 SE ranges do not overlap |
| Harvey's Legal Agent Benchmark | 25.42% ±3.67 | 8.33% ±2.00 | 17.08 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 43.75% ±3.45 | 49.04% ±3.48 | 5.29 pts | Reported ±1 SE ranges overlap |
| LegalBench | 85.26% ±0.45 | 84.84% ±0.40 | 0.42 pts | Reported ±1 SE ranges overlap |
| EMB | 56.98% ±3.13 | 56.34% ±3.32 | 0.63 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 60.60% ±0.28 | 55.84% ±2.07 | 4.76 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 56.86% ±2.32 | 73.09% ±2.97 | 16.23 pts | Reported ±1 SE ranges do not overlap |
| TaxEval v2 | 80.38% ±0.76 | 72.36% ±0.88 | 8.01 pts | Reported ±1 SE ranges do not overlap |
| MedCode | 49.35% ±2.19 | 42.86% ±2.11 | 6.48 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 90.06% ±1.96 | 88.81% ±2.00 | 1.25 pts | Reported ±1 SE ranges overlap |
| ProofBench v1.1 | 43.00% ±4.98 | 49.00% ±5.02 | 6.00 pts | Reported ±1 SE ranges overlap |
| MMLU Pro | 88.28% ±0.32 | 86.77% ±0.34 | 1.51 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 29.95% ±4.02 | 44.22% ±4.29 | 14.27 pts | Reported ±1 SE ranges do not overlap |
| IOI | 21.78% ±0.87 | 68.44% ±7.48 | 46.67 pts | Reported ±1 SE ranges do not overlap |
| ProgramBench | 0.50% ±0.50 | 1.50% ±0.86 | 1.00 pts | Reported ±1 SE ranges overlap |
| SkillsBench | 53.04% ±4.42 | 47.51% ±4.52 | 5.53 pts | Reported ±1 SE ranges overlap |
| SWE-bench | 86.60% ±1.52 | 95.40% ±0.94 | 8.80 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench 4.0 | 6.06% ±1.51 | 38.89% ±1.82 | 32.83 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 79.10% ±3.31 | 78.13% ±3.94 | 0.97 pts | Reported ±1 SE ranges overlap |
| CyberBench v1.1 | 69.52% ±5.56 | 72.08% ±5.41 | 2.56 pts | Reported ±1 SE ranges overlap |
| Public Benefits Bench v1.1 | 68.47% ±1.21 | 68.54% ±1.21 | 0.07 pts | Reported ±1 SE ranges overlap |
Performance by category
| Category | Muse Spark 1.2 average | GLM 5.3 average |
|---|---|---|
| Legal | 51.48% | 47.40% |
| Finance | 63.70% | 64.41% |
| Healthcare | 69.70% | 65.84% |
| Math | 43.00% | 49.00% |
| Academic | 88.28% | 86.77% |
| Coding | 39.58% | 53.44% |
| Cyber | 69.52% | 72.08% |
| Social Mobility | 68.47% | 68.54% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Muse Spark 1.2 cost | GLM 5.3 cost | Muse Spark 1.2 latency | GLM 5.3 latency |
|---|---|---|---|---|
| Vals Index | $1.94 | $7.25 | 22m56s | 1h13m |
| Harvey's Legal Agent Benchmark | $2.09 | $4.30 | 24m16s | 42m38s |
| Legal Research Bench | $0.51 | $2.24 | 5m07s | 51m50s |
| LegalBench | N/A | N/A | 30.10s | 22.48s |
| EMB | $2.31 | $3.79 | 17m07s | 36m12s |
| Finance Agent (v2) | $0.77 | $1.07 | 5m04s | 15m50s |
| Tax Agent Bench | $0.27 | $1.80 | 119.91s | 50m53s |
| TaxEval v2 | N/A | N/A | 48.31s | 98.32s |
| MedCode | N/A | N/A | 60.17s | 2m49s |
| MedScribe | N/A | N/A | 61.46s | 2m02s |
| ProofBench v1.1 | $0.44 | $2.08 | 7m49s | 42m23s |
| MMLU Pro | N/A | N/A | 44.23s | 61.78s |
| Code Migration | $3.78 | $24.91 | 31m07s | 3h47m |
| IOI | $2.64 | $7.67 | 25m35s | 1h46m |
| ProgramBench | $2.47 | $21.96 | 26m38s | 4h16m |
| SkillsBench | $0.65 | $0.71 | 13m00s | 13m07s |
| SWE-bench | $0.55 | $0.34 | 8m07s | 14m01s |
| Terminal-Bench 4.0 | $4.14 | $9.37 | 1h18m | 1h23m |
| Vibe Code Bench v1.1 | $1.53 | $12.45 | 19m55s | 1h04m |
| CyberBench v1.1 | $2.15 | $2.69 | 20m21s | 30m18s |
| Public Benefits Bench v1.1 | $0.59 | $0.95 | 8m55s | 35m51s |
Results available only for Muse Spark 1.2
- MortgageTax
- BioMysteryBench
- MMMU Pro
- SAGE
Results available only for GLM 5.3
- Vals RSI Index
- MysteryMechanism
- Terminal-Bench Science
- GPQA Diamond
- LiveCodeBench
- Vibe Code Bench 1-100