DeepSeek V4 Pro 0813 vs Gemini 3.5 Flash: Benchmark Comparison

DeepSeek V4 Pro 0813 has the higher score on 11 of 21 shared benchmarks; Gemini 3.5 Flash leads on 9.

The largest observed score gap is 33.61 pts on Vibe Code Bench v1.1 , where DeepSeek V4 Pro 0813 leads.

Reported ±1 standard-error ranges overlap on 9 of 21 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4 Pro 0813 Gemini 3.5 Flash Gap Reported uncertainty
Vals Index 47.63% ±1.12 44.79% ±1.10 2.84 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 7.50% ±2.15 2.50% ±0.83 5.00 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 40.87% ±3.42 30.77% ±3.21 10.10 pts Reported ±1 SE ranges do not overlap
LegalBench 82.36% ±0.44 83.60% ±0.85 1.24 pts Reported ±1 SE ranges overlap
EMB 52.80% ±3.06 63.55% ±2.75 10.75 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 50.39% ±0.25 57.86% ±0.23 7.47 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 58.66% ±3.14 53.53% ±3.30 5.14 pts Reported ±1 SE ranges overlap
TaxEval v2 73.06% ±0.87 74.37% ±0.85 1.31 pts Reported ±1 SE ranges overlap
MedCode 42.47% ±2.16 55.83% ±2.11 13.36 pts Reported ±1 SE ranges do not overlap
MedScribe 80.17% ±2.00 76.57% ±1.92 3.60 pts Reported ±1 SE ranges overlap
ProofBench v1.1 50.00% ±5.03 31.00% ±4.65 19.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 4.29% ±2.44 5.71% ±2.79 1.43 pts Reported ±1 SE ranges overlap
GPQA Diamond 92.42% ±2.02 92.68% ±1.46 0.25 pts Reported ±1 SE ranges overlap
MMLU Pro 86.97% ±0.34 89.52% ±0.31 2.54 pts Reported ±1 SE ranges do not overlap
Code Migration 41.54% ±4.30 26.75% ±4.10 14.79 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 87.53% ±0.96 87.60% ±0.95 0.08 pts Reported ±1 SE ranges overlap
ProgramBench 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
SkillsBench 53.83% ±4.49 52.74% ±4.47 1.09 pts Reported ±1 SE ranges overlap
SWE-bench 96.40% ±0.83 78.80% ±1.83 17.60 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 14.14% ±2.02 6.06% ±1.51 8.08 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 82.30% ±3.15 48.68% ±4.73 33.61 pts Reported ±1 SE ranges do not overlap

Performance by category

Category DeepSeek V4 Pro 0813 average Gemini 3.5 Flash average
Legal 43.58% 38.96%
Finance 58.73% 62.33%
Healthcare 61.32% 66.20%
Math 50.00% 31.00%
Science 4.29% 5.71%
Academic 89.70% 91.10%
Coding 53.68% 42.95%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4 Pro 0813 cost Gemini 3.5 Flash cost DeepSeek V4 Pro 0813 latency Gemini 3.5 Flash latency
Vals Index $3.35 $3.30 1h04m 18m17s
Harvey's Legal Agent Benchmark $0.17 $2.12 21m53s 7m58s
Legal Research Bench $1.09 $1.31 42m07s 4m38s
LegalBench N/A N/A 20.05s 34.58s
EMB $1.08 $5.17 37m27s 18m44s
Finance Agent (v2) $0.88 $2.51 16m15s 5m22s
Tax Agent Bench $0.73 $0.86 26m37s 3m53s
TaxEval v2 N/A N/A 3m46s 19.69s
MedCode N/A N/A 3m01s 25.29s
MedScribe N/A N/A 2m36s 57.08s
ProofBench v1.1 $0.06 $0.57 12m12s 6m34s
Terminal-Bench Science $4.63 $3.45 2h16m 59m32s
GPQA Diamond N/A N/A 4m39s 27.66s
MMLU Pro N/A N/A 59.62s 26.26s
Code Migration $18.58 $5.73 3h32m 19m28s
LiveCodeBench N/A N/A 4m33s 64.37s
ProgramBench N/A N/A 1h23m 18m25s
SkillsBench $0.37 $1.25 11m15s 7m15s
SWE-bench $0.10 $0.95 4m00s 4m14s
Terminal-Bench 4.0 $3.31 $5.98 1h32m 1h11m
Vibe Code Bench v1.1 $0.36 $2.54 1h06m 14m53s

Results available only for DeepSeek V4 Pro 0813

  • IOI
  • Vibe Code Bench 1-100
  • CyberBench v1.1

Results available only for Gemini 3.5 Flash

  • MortgageTax
  • MMMU Pro
  • SAGE
  • Public Benefits Bench v1.1
Model details DeepSeek V4 Pro 0813 Model details Gemini 3.5 Flash