DeepSeek V4.1 Flash vs GPT 5.5: Benchmark Comparison

DeepSeek V4.1 Flash has the higher score on 8 of 15 shared benchmarks; GPT 5.5 leads on 6.

The largest observed score gap is 14.89 pts on Vibe Code Bench v1.1 , where DeepSeek V4.1 Flash leads.

Reported ±1 standard-error ranges overlap on 8 of 15 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4.1 Flash GPT 5.5 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 6.67% ±1.83 3.75% ±1.17 2.92 pts Reported ±1 SE ranges overlap
Legal Research Bench 41.35% ±3.42 40.38% ±3.41 0.96 pts Reported ±1 SE ranges overlap
LegalBench 83.28% ±0.46 86.52% ±0.41 3.24 pts Reported ±1 SE ranges do not overlap
EMB 57.21% ±3.23 64.54% ±2.87 7.34 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 53.48% ±0.39 51.76% ±0.55 1.72 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.46% ±3.15 60.46% ±3.22 2.00 pts Reported ±1 SE ranges overlap
MedCode 41.17% ±2.04 49.10% ±2.19 7.93 pts Reported ±1 SE ranges do not overlap
MedScribe 85.50% ±1.92 86.87% ±1.93 1.37 pts Reported ±1 SE ranges overlap
SAGE 47.88% ±3.44 51.53% ±3.95 3.65 pts Reported ±1 SE ranges overlap
Code Migration 45.62% ±4.29 45.16% ±4.16 0.46 pts Reported ±1 SE ranges overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
SkillsBench 69.80% ±3.88 62.21% ±4.52 7.59 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 84.74% ±2.89 69.85% ±4.54 14.89 pts Reported ±1 SE ranges do not overlap
SRE Bench 0.76% ±0.54 3.82% ±1.19 3.05 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 64.28% ±1.25 60.89% ±1.27 3.38 pts Reported ±1 SE ranges do not overlap

Performance by category

Category DeepSeek V4.1 Flash average GPT 5.5 average
Legal 43.76% 43.55%
Finance 57.72% 58.92%
Healthcare 63.34% 67.98%
Education 47.88% 51.53%
Coding 50.17% 44.43%
Cyber 0.76% 3.82%
Social Mobility 64.28% 60.89%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4.1 Flash cost GPT 5.5 cost DeepSeek V4.1 Flash latency GPT 5.5 latency
Harvey's Legal Agent Benchmark $0.17 $4.60 9m39s 12m14s
Legal Research Bench $0.25 $7.40 11m42s 37m45s
LegalBench N/A N/A 6.53s 18.14s
EMB $0.23 $3.27 13m00s 15m11s
Finance Agent (v2) $0.21 $4.15 6m04s 11m02s
Tax Agent Bench $0.13 $4.07 5m49s 15m52s
MedCode N/A N/A 35.13s 2m40s
MedScribe N/A N/A 53.83s 2m13s
SAGE N/A N/A 35.78s 76.05s
Code Migration $0.94 $6.44 1h45m 31m35s
ProgramBench $0.90 $6.95 1h02m 22m03s
SkillsBench $0.09 $2.54 3m43s 6m55s
Vibe Code Bench v1.1 $0.41 $16.66 15m28s 31m52s
SRE Bench $0.55 $17.77 1h01m 46m28s
Public Benefits Bench v1.1 $0.07 $3.97 9m59s 39m37s

Results available only for DeepSeek V4.1 Flash

  • Vals Index
  • ProofBench v1.1
  • BioMysteryBench
  • MysteryMechanism
  • Terminal-Bench Science
  • IOI
  • Terminal-Bench 4.0
  • Vibe Code Bench 1-100
  • CyberBench v1.1

Results available only for GPT 5.5

  • Vals RSI Index
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SWE-bench
  • Time Horizon Index: KSP
Model details DeepSeek V4.1 Flash Model details GPT 5.5