Hy4 Preview vs GLM 5.3: Benchmark Comparison

Hy4 Preview has the higher score on 6 of 18 shared benchmarks; GLM 5.3 leads on 12.

The largest observed score gap is 30.81 pts on Terminal-Bench 4.0 , where GLM 5.3 leads.

Reported ±1 standard-error ranges overlap on 10 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Hy4 Preview GLM 5.3 Gap Reported uncertainty
Vals Index 49.94% ±1.15 53.51% ±1.30 3.57 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 9.17% ±2.00 8.33% ±2.00 0.83 pts Reported ±1 SE ranges overlap
Legal Research Bench 45.19% ±3.46 49.04% ±3.48 3.85 pts Reported ±1 SE ranges overlap
LegalBench 83.76% ±0.41 84.84% ±0.40 1.08 pts Reported ±1 SE ranges do not overlap
EMB 58.08% ±3.14 56.34% ±3.32 1.73 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 55.06% ±0.31 55.84% ±2.07 0.78 pts Reported ±1 SE ranges overlap
Tax Agent Bench 63.71% ±3.26 73.09% ±2.97 9.38 pts Reported ±1 SE ranges do not overlap
MedCode 43.25% ±2.13 42.86% ±2.11 0.38 pts Reported ±1 SE ranges overlap
MedScribe 83.60% ±2.06 88.81% ±2.00 5.21 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 75.00% ±4.35 49.00% ±5.02 26.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 1.43% ±1.43 5.71% ±2.79 4.29 pts Reported ±1 SE ranges do not overlap
Code Migration 47.43% ±4.27 44.22% ±4.29 3.21 pts Reported ±1 SE ranges overlap
IOI 59.33% ±4.60 68.44% ±7.48 9.11 pts Reported ±1 SE ranges overlap
ProgramBench 0.00% ±0.00 1.50% ±0.86 1.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 8.08% ±1.34 38.89% ±1.82 30.81 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 77.48% ±4.04 78.13% ±3.94 0.65 pts Reported ±1 SE ranges overlap
CyberBench v1.1 65.36% ±5.55 72.08% ±5.41 6.73 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 68.61% ±1.21 68.54% ±1.21 0.07 pts Reported ±1 SE ranges overlap

Performance by category

Category Hy4 Preview average GLM 5.3 average
Legal 46.04% 47.40%
Finance 58.95% 61.76%
Healthcare 63.42% 65.84%
Math 75.00% 49.00%
Science 1.43% 5.71%
Coding 38.47% 46.24%
Cyber 65.36% 72.08%
Social Mobility 68.61% 68.54%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Hy4 Preview cost GLM 5.3 cost Hy4 Preview latency GLM 5.3 latency
Vals Index $1.44 $7.25 1h01m 1h13m
Harvey's Legal Agent Benchmark $0.98 $4.30 29m25s 42m38s
Legal Research Bench $0.85 $2.24 1h04m 51m50s
LegalBench N/A N/A 90.56s 22.48s
EMB $0.79 $3.79 30m39s 36m12s
Finance Agent (v2) $0.59 $1.07 23m45s 15m50s
Tax Agent Bench $0.59 $1.80 41m04s 50m53s
MedCode N/A N/A 6m39s 2m49s
MedScribe N/A N/A 5m50s 2m02s
ProofBench v1.1 $0.34 $2.08 22m54s 42m23s
Terminal-Bench Science $2.14 $15.63 3h48m 2h45m
Code Migration $3.41 $24.91 2h13m 3h47m
IOI $1.35 $7.67 1h01m 1h46m
ProgramBench $12.81 $21.96 2h10m 4h16m
Terminal-Bench 4.0 $2.10 $9.37 2h10m 1h23m
Vibe Code Bench v1.1 $2.18 $12.45 41m33s 1h04m
CyberBench v1.1 $0.49 $2.69 36m43s 30m18s
Public Benefits Bench v1.1 $0.25 $0.95 1h03m 35m51s

Results available only for Hy4 Preview

  • BioMysteryBench
  • SRE Bench

Results available only for GLM 5.3

  • Vals RSI Index
  • TaxEval v2
  • MysteryMechanism
  • GPQA Diamond
  • MMLU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
Model details Hy4 Preview Model details GLM 5.3