DeepSeek V4.1 Flash vs Hy4 Preview: Benchmark Comparison

DeepSeek V4.1 Flash has the higher score on 6 of 20 shared benchmarks; Hy4 Preview leads on 14.

The largest observed score gap is 21.00 pts on ProofBench v1.1 , where Hy4 Preview leads.

Reported ±1 standard-error ranges overlap on 13 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4.1 Flash Hy4 Preview Gap Reported uncertainty
Vals Index 51.32% ±1.13 49.94% ±1.15 1.38 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 6.67% ±1.83 9.17% ±2.00 2.50 pts Reported ±1 SE ranges overlap
Legal Research Bench 41.35% ±3.42 45.19% ±3.46 3.85 pts Reported ±1 SE ranges overlap
LegalBench 83.28% ±0.46 83.76% ±0.41 0.48 pts Reported ±1 SE ranges overlap
EMB 57.21% ±3.23 58.08% ±3.14 0.87 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.48% ±0.39 55.06% ±0.31 1.58 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.46% ±3.15 63.71% ±3.26 1.25 pts Reported ±1 SE ranges overlap
MedCode 41.17% ±2.04 43.25% ±2.13 2.07 pts Reported ±1 SE ranges overlap
MedScribe 85.50% ±1.92 83.60% ±2.06 1.90 pts Reported ±1 SE ranges overlap
ProofBench v1.1 54.00% ±5.01 75.00% ±4.35 21.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 67.78% ±1.11 69.26% ±2.59 1.48 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 0.00% ±0.00 1.43% ±1.43 1.43 pts Reported ±1 SE ranges overlap
Code Migration 45.62% ±4.29 47.43% ±4.27 1.81 pts Reported ±1 SE ranges overlap
IOI 40.28% ±2.56 59.33% ±4.60 19.05 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.00% ±0.00 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 19.70% ±1.75 8.08% ±1.34 11.62 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 84.74% ±2.89 77.48% ±4.04 7.26 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 73.69% ±5.48 65.36% ±5.55 8.33 pts Reported ±1 SE ranges overlap
SRE Bench 0.76% ±0.54 2.29% ±0.93 1.53 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 64.28% ±1.25 68.61% ±1.21 4.33 pts Reported ±1 SE ranges do not overlap

Performance by category

Category DeepSeek V4.1 Flash average Hy4 Preview average
Legal 43.76% 46.04%
Finance 57.72% 58.95%
Healthcare 63.34% 63.42%
Math 54.00% 75.00%
Science 33.89% 35.34%
Coding 38.17% 38.47%
Cyber 37.23% 33.82%
Social Mobility 64.28% 68.61%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4.1 Flash cost Hy4 Preview cost DeepSeek V4.1 Flash latency Hy4 Preview latency
Vals Index $0.33 $1.44 27m53s 1h01m
Harvey's Legal Agent Benchmark $0.17 $0.98 9m39s 29m25s
Legal Research Bench $0.25 $0.85 11m42s 1h04m
LegalBench N/A N/A 6.53s 90.56s
EMB $0.23 $0.79 13m00s 30m39s
Finance Agent (v2) $0.21 $0.59 6m04s 23m45s
Tax Agent Bench $0.13 $0.59 5m49s 41m04s
MedCode N/A N/A 35.13s 6m39s
MedScribe N/A N/A 53.83s 5m50s
ProofBench v1.1 $0.13 $0.34 9m07s 22m54s
BioMysteryBench $0.23 $0.36 14m50s 23m30s
Terminal-Bench Science $0.47 $2.14 2h57m 3h48m
Code Migration $0.94 $3.41 1h45m 2h13m
IOI $0.24 $1.35 18m30s 1h01m
ProgramBench $0.90 $12.81 1h02m 2h10m
Terminal-Bench 4.0 $0.50 $2.10 48m22s 2h10m
Vibe Code Bench v1.1 $0.41 $2.18 15m28s 41m33s
CyberBench v1.1 $0.07 $0.49 13m50s 36m43s
SRE Bench $0.55 $4.90 1h01m 1h40m
Public Benefits Bench v1.1 $0.07 $0.25 9m59s 1h03m

Results available only for DeepSeek V4.1 Flash

  • MysteryMechanism
  • SAGE
  • SkillsBench
  • Vibe Code Bench 1-100

Results available only for Hy4 Preview

None.

Model details DeepSeek V4.1 Flash Model details Hy4 Preview