DeepSeek V4.1 Flash vs Gemini 3.7 Flash: Benchmark Comparison

DeepSeek V4.1 Flash has the higher score on 10 of 20 shared benchmarks; Gemini 3.7 Flash leads on 10.

The largest observed score gap is 29.94 pts on CyberBench v1.1 , where DeepSeek V4.1 Flash leads.

Reported ±1 standard-error ranges overlap on 8 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4.1 Flash Gemini 3.7 Flash Gap Reported uncertainty
Vals Index 51.32% ±1.13 51.27% ±1.10 0.05 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 6.67% ±1.83 8.75% ±2.15 2.08 pts Reported ±1 SE ranges overlap
Legal Research Bench 41.35% ±3.42 34.62% ±3.31 6.73 pts Reported ±1 SE ranges do not overlap
LegalBench 83.28% ±0.46 87.26% ±0.42 3.98 pts Reported ±1 SE ranges do not overlap
EMB 57.21% ±3.23 71.33% ±2.26 14.12 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 53.48% ±0.39 59.04% ±0.27 5.56 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.46% ±3.15 57.66% ±3.31 4.79 pts Reported ±1 SE ranges overlap
MedCode 41.17% ±2.04 53.39% ±2.12 12.22 pts Reported ±1 SE ranges do not overlap
MedScribe 85.50% ±1.92 83.94% ±2.00 1.56 pts Reported ±1 SE ranges overlap
ProofBench v1.1 54.00% ±5.01 58.00% ±4.96 4.00 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 0.00% ±0.00 7.14% ±3.10 7.14 pts Reported ±1 SE ranges do not overlap
SAGE 47.88% ±3.44 49.23% ±3.38 1.35 pts Reported ±1 SE ranges overlap
Code Migration 45.62% ±4.29 34.80% ±4.22 10.82 pts Reported ±1 SE ranges do not overlap
IOI 40.28% ±2.56 67.83% ±3.61 27.55 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.00% ±0.00 0.50 pts Reported ±1 SE ranges overlap
SkillsBench 69.80% ±3.88 65.89% ±4.55 3.91 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 19.70% ±1.75 12.12% ±3.03 7.58 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 84.74% ±2.89 70.39% ±4.84 14.34 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 73.69% ±5.48 43.75% ±2.21 29.94 pts Reported ±1 SE ranges do not overlap
SRE Bench 0.76% ±0.54 4.58% ±1.29 3.82 pts Reported ±1 SE ranges do not overlap

Performance by category

Category DeepSeek V4.1 Flash average Gemini 3.7 Flash average
Legal 43.76% 43.54%
Finance 57.72% 62.68%
Healthcare 63.34% 68.67%
Math 54.00% 58.00%
Science 0.00% 7.14%
Education 47.88% 49.23%
Coding 43.44% 41.84%
Cyber 37.23% 24.16%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4.1 Flash cost Gemini 3.7 Flash cost DeepSeek V4.1 Flash latency Gemini 3.7 Flash latency
Vals Index $0.33 $4.71 27m53s 47m27s
Harvey's Legal Agent Benchmark $0.17 $2.56 9m39s 7m32s
Legal Research Bench $0.25 $0.76 11m42s 4m49s
LegalBench N/A N/A 6.53s 2.23s
EMB $0.23 $5.69 13m00s 10m51s
Finance Agent (v2) $0.21 $1.48 6m04s 3m01s
Tax Agent Bench $0.13 $0.48 5m49s 87.24s
MedCode N/A N/A 35.13s 9.16s
MedScribe N/A N/A 53.83s 16.82s
ProofBench v1.1 $0.13 $0.56 9m07s 7m04s
Terminal-Bench Science $0.47 $6.26 2h57m 1h18m
SAGE N/A N/A 35.78s 9.79s
Code Migration $0.94 $21.46 1h45m 2h39m
IOI $0.24 $3.40 18m30s 21m09s
ProgramBench $0.90 $7.13 1h02m 21m30s
SkillsBench $0.09 $1.80 3m43s 3m06s
Terminal-Bench 4.0 $0.50 $9.02 48m22s 1h24m
Vibe Code Bench v1.1 $0.41 $4.83 15m28s 23m43s
CyberBench v1.1 $0.07 $2.15 13m50s 7m34s
SRE Bench $0.55 $21.47 1h01m 1h50m

Results available only for DeepSeek V4.1 Flash

  • BioMysteryBench
  • MysteryMechanism
  • Vibe Code Bench 1-100
  • Public Benefits Bench v1.1

Results available only for Gemini 3.7 Flash

  • Vals RSI Index
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SWE-bench
Model details DeepSeek V4.1 Flash Model details Gemini 3.7 Flash