DeepSeek V4 Pro 0813 vs Gemini 3.6 Flash: Benchmark Comparison

DeepSeek V4 Pro 0813 has the higher score on 9 of 19 shared benchmarks; Gemini 3.6 Flash leads on 8.

The largest observed score gap is 27.50 pts on CyberBench v1.1 , where DeepSeek V4 Pro 0813 leads.

Reported ±1 standard-error ranges overlap on 5 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4 Pro 0813 Gemini 3.6 Flash Gap Reported uncertainty
Harvey's Legal Agent Benchmark 7.50% ±2.15 3.33% ±1.17 4.17 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 40.87% ±3.42 25.00% ±3.01 15.87 pts Reported ±1 SE ranges do not overlap
LegalBench 82.36% ±0.44 86.70% ±0.41 4.34 pts Reported ±1 SE ranges do not overlap
EMB 52.80% ±3.06 65.41% ±2.65 12.61 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 50.39% ±0.25 56.30% ±0.18 5.90 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 58.66% ±3.14 49.90% ±3.30 8.77 pts Reported ±1 SE ranges do not overlap
TaxEval v2 73.06% ±0.87 74.86% ±0.85 1.80 pts Reported ±1 SE ranges do not overlap
MedCode 42.47% ±2.16 53.15% ±2.16 10.68 pts Reported ±1 SE ranges do not overlap
MedScribe 80.17% ±2.00 79.66% ±1.86 0.51 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 4.29% ±2.44 4.29% ±2.44 0.00 pts Reported ±1 SE ranges overlap
GPQA Diamond 92.42% ±2.02 93.43% ±1.33 1.01 pts Reported ±1 SE ranges overlap
MMLU Pro 86.97% ±0.34 89.28% ±0.30 2.30 pts Reported ±1 SE ranges do not overlap
Code Migration 41.54% ±4.30 30.93% ±4.07 10.61 pts Reported ±1 SE ranges do not overlap
IOI 51.61% ±2.40 35.06% ±5.70 16.55 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 87.53% ±0.96 88.08% ±0.94 0.55 pts Reported ±1 SE ranges overlap
ProgramBench 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
SWE-bench 96.40% ±0.83 79.60% ±1.80 16.80 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 82.30% ±3.15 64.00% ±4.29 18.29 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 72.86% ±5.50 45.36% ±3.75 27.50 pts Reported ±1 SE ranges do not overlap

Performance by category

Category DeepSeek V4 Pro 0813 average Gemini 3.6 Flash average
Legal 43.58% 38.35%
Finance 58.73% 61.61%
Healthcare 61.32% 66.41%
Science 4.29% 4.29%
Academic 89.70% 91.35%
Coding 59.90% 49.61%
Cyber 72.86% 45.36%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4 Pro 0813 cost Gemini 3.6 Flash cost DeepSeek V4 Pro 0813 latency Gemini 3.6 Flash latency
Harvey's Legal Agent Benchmark $0.17 $1.90 21m53s 16m24s
Legal Research Bench $1.09 $0.47 42m07s 2m48s
LegalBench N/A N/A 20.05s 2.96s
EMB $1.08 $3.89 37m27s 12m06s
Finance Agent (v2) $0.88 $1.40 16m15s 3m45s
Tax Agent Bench $0.73 $0.36 26m37s 105.32s
TaxEval v2 N/A N/A 3m46s 12.34s
MedCode N/A N/A 3m01s 16.68s
MedScribe N/A N/A 2m36s 33.55s
Terminal-Bench Science $4.63 $4.39 2h16m 1h52m
GPQA Diamond N/A N/A 4m39s 18.06s
MMLU Pro N/A N/A 59.62s 7.30s
Code Migration $18.58 $8.93 3h32m 1h30m
IOI $2.15 $8.67 1h07m 48m23s
LiveCodeBench N/A N/A 4m33s 31.88s
ProgramBench $0.43 $5.92 1h23m 47m33s
SWE-bench $0.10 $1.19 4m00s 13m08s
Vibe Code Bench v1.1 $0.36 $3.04 1h06m 25m23s
CyberBench v1.1 $0.59 $1.31 22m33s 11m00s

Results available only for DeepSeek V4 Pro 0813

  • Vals Index
  • ProofBench v1.1
  • SkillsBench
  • Terminal-Bench 4.0
  • Vibe Code Bench 1-100

Results available only for Gemini 3.6 Flash

  • MortgageTax
  • BioMysteryBench
  • MMMU Pro
  • SAGE
  • Public Benefits Bench v1.1
Model details DeepSeek V4 Pro 0813 Model details Gemini 3.6 Flash