GPT 5.5 vs Hy4 Preview: Benchmark Comparison
GPT 5.5 has the higher score on 6 of 13 shared benchmarks; Hy4 Preview leads on 7.
The largest observed score gap is 7.71 pts on Public Benefits Bench v1.1 , where Hy4 Preview leads.
Reported ±1 standard-error ranges overlap on 7 of 13 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | GPT 5.5 | Hy4 Preview | Gap | Reported uncertainty |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | 3.75% ±1.17 | 9.17% ±2.00 | 5.42 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 40.38% ±3.41 | 45.19% ±3.46 | 4.81 pts | Reported ±1 SE ranges overlap |
| LegalBench | 86.52% ±0.41 | 83.76% ±0.41 | 2.76 pts | Reported ±1 SE ranges do not overlap |
| EMB | 64.54% ±2.87 | 58.08% ±3.14 | 6.47 pts | Reported ±1 SE ranges do not overlap |
| Finance Agent (v2) | 51.76% ±0.55 | 55.06% ±0.31 | 3.30 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 60.46% ±3.22 | 63.71% ±3.26 | 3.26 pts | Reported ±1 SE ranges overlap |
| MedCode | 49.10% ±2.19 | 43.25% ±2.13 | 5.85 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 86.87% ±1.93 | 83.60% ±2.06 | 3.27 pts | Reported ±1 SE ranges overlap |
| Code Migration | 45.16% ±4.16 | 47.43% ±4.27 | 2.27 pts | Reported ±1 SE ranges overlap |
| ProgramBench | 0.50% ±0.50 | 0.00% ±0.00 | 0.50 pts | Reported ±1 SE ranges overlap |
| Vibe Code Bench v1.1 | 69.85% ±4.54 | 77.48% ±4.04 | 7.63 pts | Reported ±1 SE ranges overlap |
| SRE Bench | 3.82% ±1.19 | 2.29% ±0.93 | 1.53 pts | Reported ±1 SE ranges overlap |
| Public Benefits Bench v1.1 | 60.89% ±1.27 | 68.61% ±1.21 | 7.71 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | GPT 5.5 average | Hy4 Preview average |
|---|---|---|
| Legal | 43.55% | 46.04% |
| Finance | 58.92% | 58.95% |
| Healthcare | 67.98% | 63.42% |
| Coding | 38.50% | 41.64% |
| Cyber | 3.82% | 2.29% |
| Social Mobility | 60.89% | 68.61% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | GPT 5.5 cost | Hy4 Preview cost | GPT 5.5 latency | Hy4 Preview latency |
|---|---|---|---|---|
| Harvey's Legal Agent Benchmark | $4.60 | $0.98 | 12m14s | 29m25s |
| Legal Research Bench | $7.40 | $0.85 | 37m45s | 1h04m |
| LegalBench | N/A | N/A | 18.14s | 90.56s |
| EMB | $3.27 | $0.79 | 15m11s | 30m39s |
| Finance Agent (v2) | $4.15 | $0.59 | 11m02s | 23m45s |
| Tax Agent Bench | $4.07 | $0.59 | 15m52s | 41m04s |
| MedCode | N/A | N/A | 2m40s | 6m39s |
| MedScribe | N/A | N/A | 2m13s | 5m50s |
| Code Migration | $6.44 | $3.41 | 31m35s | 2h13m |
| ProgramBench | $6.95 | $12.81 | 22m03s | 2h10m |
| Vibe Code Bench v1.1 | $16.66 | $2.18 | 31m52s | 41m33s |
| SRE Bench | $17.77 | $4.90 | 46m28s | 1h40m |
| Public Benefits Bench v1.1 | $3.97 | $0.25 | 39m37s | 1h03m |
Results available only for GPT 5.5
- Vals RSI Index
- MortgageTax
- TaxEval v2
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- SAGE
- LiveCodeBench
- SkillsBench
- SWE-bench
- Time Horizon Index: KSP
Results available only for Hy4 Preview
- Vals Index
- ProofBench v1.1
- BioMysteryBench
- Terminal-Bench Science
- IOI
- Terminal-Bench 4.0
- CyberBench v1.1