Muse Spark 1.2 vs Hy4 Preview: Benchmark Comparison

Muse Spark 1.2 has the higher score on 8 of 18 shared benchmarks; Hy4 Preview leads on 10.

The largest observed score gap is 37.55 pts on IOI , where Hy4 Preview leads.

Reported ±1 standard-error ranges overlap on 8 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Muse Spark 1.2 Hy4 Preview Gap Reported uncertainty
Vals Index 49.29% ±1.10 49.94% ±1.15 0.66 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 25.42% ±3.67 9.17% ±2.00 16.25 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 43.75% ±3.45 45.19% ±3.46 1.44 pts Reported ±1 SE ranges overlap
LegalBench 85.26% ±0.45 83.76% ±0.41 1.50 pts Reported ±1 SE ranges do not overlap
EMB 56.98% ±3.13 58.08% ±3.14 1.10 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 60.60% ±0.28 55.06% ±0.31 5.54 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 56.86% ±2.32 63.71% ±3.26 6.85 pts Reported ±1 SE ranges do not overlap
MedCode 49.35% ±2.19 43.25% ±2.13 6.10 pts Reported ±1 SE ranges do not overlap
MedScribe 90.06% ±1.96 83.60% ±2.06 6.46 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 43.00% ±4.98 75.00% ±4.35 32.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 64.81% ±1.33 69.26% ±2.59 4.44 pts Reported ±1 SE ranges do not overlap
Code Migration 29.95% ±4.02 47.43% ±4.27 17.48 pts Reported ±1 SE ranges do not overlap
IOI 21.78% ±0.87 59.33% ±4.60 37.55 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.00% ±0.00 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 6.06% ±1.51 8.08% ±1.34 2.02 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 79.10% ±3.31 77.48% ±4.04 1.62 pts Reported ±1 SE ranges overlap
CyberBench v1.1 69.52% ±5.56 65.36% ±5.55 4.17 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 68.47% ±1.21 68.61% ±1.21 0.13 pts Reported ±1 SE ranges overlap

Performance by category

Category Muse Spark 1.2 average Hy4 Preview average
Legal 51.48% 46.04%
Finance 58.15% 58.95%
Healthcare 69.70% 63.42%
Math 43.00% 75.00%
Science 64.81% 69.26%
Coding 27.48% 38.47%
Cyber 69.52% 65.36%
Social Mobility 68.47% 68.61%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Muse Spark 1.2 cost Hy4 Preview cost Muse Spark 1.2 latency Hy4 Preview latency
Vals Index $1.94 $1.44 22m56s 1h01m
Harvey's Legal Agent Benchmark $2.09 $0.98 24m16s 29m25s
Legal Research Bench $0.51 $0.85 5m07s 1h04m
LegalBench N/A N/A 30.10s 90.56s
EMB $2.31 $0.79 17m07s 30m39s
Finance Agent (v2) $0.77 $0.59 5m04s 23m45s
Tax Agent Bench $0.27 $0.59 119.91s 41m04s
MedCode N/A N/A 60.17s 6m39s
MedScribe N/A N/A 61.46s 5m50s
ProofBench v1.1 $0.44 $0.34 7m49s 22m54s
BioMysteryBench $0.99 $0.36 23m37s 23m30s
Code Migration $3.78 $3.41 31m07s 2h13m
IOI $2.64 $1.35 25m35s 1h01m
ProgramBench $2.47 $12.81 26m38s 2h10m
Terminal-Bench 4.0 $4.14 $2.10 1h18m 2h10m
Vibe Code Bench v1.1 $1.53 $2.18 19m55s 41m33s
CyberBench v1.1 $2.15 $0.49 20m21s 36m43s
Public Benefits Bench v1.1 $0.59 $0.25 8m55s 1h03m

Results available only for Muse Spark 1.2

  • MortgageTax
  • TaxEval v2
  • MMLU Pro
  • MMMU Pro
  • SAGE
  • SkillsBench
  • SWE-bench

Results available only for Hy4 Preview

  • Terminal-Bench Science
  • SRE Bench
Model details Muse Spark 1.2 Model details Hy4 Preview