DeepSeek V4 Pro 0813 vs Hy4 Preview: Benchmark Comparison

DeepSeek V4 Pro 0813 has the higher score on 4 of 17 shared benchmarks; Hy4 Preview leads on 12.

The largest observed score gap is 25.00 pts on ProofBench v1.1 , where Hy4 Preview leads.

Reported ±1 standard-error ranges overlap on 11 of 17 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4 Pro 0813 Hy4 Preview Gap Reported uncertainty
Vals Index 47.63% ±1.12 49.94% ±1.15 2.31 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 7.50% ±2.15 9.17% ±2.00 1.67 pts Reported ±1 SE ranges overlap
Legal Research Bench 40.87% ±3.42 45.19% ±3.46 4.33 pts Reported ±1 SE ranges overlap
LegalBench 82.36% ±0.44 83.76% ±0.41 1.40 pts Reported ±1 SE ranges do not overlap
EMB 52.80% ±3.06 58.08% ±3.14 5.27 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 50.39% ±0.25 55.06% ±0.31 4.67 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 58.66% ±3.14 63.71% ±3.26 5.05 pts Reported ±1 SE ranges overlap
MedCode 42.47% ±2.16 43.25% ±2.13 0.78 pts Reported ±1 SE ranges overlap
MedScribe 80.17% ±2.00 83.60% ±2.06 3.43 pts Reported ±1 SE ranges overlap
ProofBench v1.1 50.00% ±5.03 75.00% ±4.35 25.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 4.29% ±2.44 1.43% ±1.43 2.86 pts Reported ±1 SE ranges overlap
Code Migration 41.54% ±4.30 47.43% ±4.27 5.89 pts Reported ±1 SE ranges overlap
IOI 51.61% ±2.40 59.33% ±4.60 7.72 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 14.14% ±2.02 8.08% ±1.34 6.06 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 82.30% ±3.15 77.48% ±4.04 4.82 pts Reported ±1 SE ranges overlap
CyberBench v1.1 72.86% ±5.50 65.36% ±5.55 7.50 pts Reported ±1 SE ranges overlap

Performance by category

Category DeepSeek V4 Pro 0813 average Hy4 Preview average
Legal 43.58% 46.04%
Finance 53.95% 58.95%
Healthcare 61.32% 63.42%
Math 50.00% 75.00%
Science 4.29% 1.43%
Coding 37.92% 38.47%
Cyber 72.86% 65.36%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4 Pro 0813 cost Hy4 Preview cost DeepSeek V4 Pro 0813 latency Hy4 Preview latency
Vals Index $3.35 $1.44 1h04m 1h01m
Harvey's Legal Agent Benchmark $0.17 $0.98 21m53s 29m25s
Legal Research Bench $1.09 $0.85 42m07s 1h04m
LegalBench N/A N/A 20.05s 90.56s
EMB $1.08 $0.79 37m27s 30m39s
Finance Agent (v2) $0.88 $0.59 16m15s 23m45s
Tax Agent Bench $0.73 $0.59 26m37s 41m04s
MedCode N/A N/A 3m01s 6m39s
MedScribe N/A N/A 2m36s 5m50s
ProofBench v1.1 $0.06 $0.34 12m12s 22m54s
Terminal-Bench Science $4.63 $2.14 2h16m 3h48m
Code Migration $18.58 $3.41 3h32m 2h13m
IOI $2.15 $1.35 1h07m 1h01m
ProgramBench $0.43 $12.81 1h23m 2h10m
Terminal-Bench 4.0 $3.31 $2.10 1h32m 2h10m
Vibe Code Bench v1.1 $0.36 $2.18 1h06m 41m33s
CyberBench v1.1 $0.59 $0.49 22m33s 36m43s

Results available only for DeepSeek V4 Pro 0813

  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100

Results available only for Hy4 Preview

  • BioMysteryBench
  • SRE Bench
  • Public Benefits Bench v1.1
Model details DeepSeek V4 Pro 0813 Model details Hy4 Preview