Hy4 Preview vs GLM 5.2: Benchmark Comparison
Hy4 Preview has the higher score on 8 of 11 shared benchmarks; GLM 5.2 leads on 3.
The largest observed score gap is 13.94 pts on Legal Research Bench , where Hy4 Preview leads.
Reported ±1 standard-error ranges overlap on 6 of 11 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Hy4 Preview | GLM 5.2 | Gap | Reported uncertainty |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | 9.17% ±2.00 | 7.08% ±2.00 | 2.08 pts | Reported ±1 SE ranges overlap |
| Legal Research Bench | 45.19% ±3.46 | 31.25% ±3.22 | 13.94 pts | Reported ±1 SE ranges do not overlap |
| LegalBench | 83.76% ±0.41 | 84.07% ±0.45 | 0.32 pts | Reported ±1 SE ranges overlap |
| EMB | 58.08% ±3.14 | 61.53% ±2.88 | 3.46 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 55.06% ±0.31 | 49.70% ±0.88 | 5.36 pts | Reported ±1 SE ranges do not overlap |
| MedCode | 43.25% ±2.13 | 40.77% ±2.17 | 2.48 pts | Reported ±1 SE ranges overlap |
| MedScribe | 83.60% ±2.06 | 83.53% ±2.00 | 0.07 pts | Reported ±1 SE ranges overlap |
| Code Migration | 47.43% ±4.27 | 37.87% ±4.14 | 9.56 pts | Reported ±1 SE ranges do not overlap |
| ProgramBench | 0.00% ±0.00 | 0.50% ±0.50 | 0.50 pts | Reported ±1 SE ranges overlap |
| Vibe Code Bench v1.1 | 77.48% ±4.04 | 63.96% ±4.79 | 13.52 pts | Reported ±1 SE ranges do not overlap |
| SRE Bench | 2.29% ±0.93 | 0.00% ±0.00 | 2.29 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Hy4 Preview average | GLM 5.2 average |
|---|---|---|
| Legal | 46.04% | 40.80% |
| Finance | 56.57% | 55.62% |
| Healthcare | 63.42% | 62.15% |
| Coding | 41.64% | 34.11% |
| Cyber | 2.29% | 0.00% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Hy4 Preview cost | GLM 5.2 cost | Hy4 Preview latency | GLM 5.2 latency |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | $0.98 | $2.06 | 29m25s | 20m51s |
| Legal Research Bench | $0.85 | $0.89 | 1h04m | 17m03s |
| LegalBench | N/A | N/A | 90.56s | 5.77s |
| EMB | $0.79 | $4.47 | 30m39s | 58m52s |
| Finance Agent (v2) | $0.59 | $0.71 | 23m45s | 7m53s |
| MedCode | N/A | N/A | 6m39s | 95.27s |
| MedScribe | N/A | N/A | 5m50s | 2m18s |
| Code Migration | $3.41 | $12.64 | 2h13m | 1h54m |
| ProgramBench | $12.81 | $12.58 | 2h10m | 2h06m |
| Vibe Code Bench v1.1 | $2.18 | $8.49 | 41m33s | 1h02m |
| SRE Bench | $4.90 | $26.58 | 1h40m | 2h29m |
Results available only for Hy4 Preview
- Vals Index
- Tax Agent Bench
- ProofBench v1.1
- BioMysteryBench
- Terminal-Bench Science
- IOI
- Terminal-Bench 4.0
- CyberBench v1.1
- Public Benefits Bench v1.1
Results available only for GLM 5.2
- TaxEval v2
- GPQA Diamond
- MMLU Pro
- LiveCodeBench
- SkillsBench
- SWE-bench