DeepSeek V4 Pro 0813 vs Grok 4.5: Benchmark Comparison

DeepSeek V4 Pro 0813 has the higher score on 12 of 22 shared benchmarks; Grok 4.5 leads on 9.

The largest observed score gap is 19.00 pts on ProofBench v1.1 , where DeepSeek V4 Pro 0813 leads.

Reported ±1 standard-error ranges overlap on 10 of 22 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4 Pro 0813 Grok 4.5 Gap Reported uncertainty
Vals Index 47.63% ±1.12 44.72% ±1.17 2.91 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 7.50% ±2.15 12.92% ±2.53 5.42 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 40.87% ±3.42 37.98% ±3.37 2.88 pts Reported ±1 SE ranges overlap
LegalBench 82.36% ±0.44 85.97% ±0.41 3.60 pts Reported ±1 SE ranges do not overlap
EMB 52.80% ±3.06 52.91% ±3.11 0.11 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 50.39% ±0.25 48.35% ±0.46 2.04 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 58.66% ±3.14 64.32% ±3.16 5.66 pts Reported ±1 SE ranges overlap
TaxEval v2 73.06% ±0.87 71.67% ±0.89 1.39 pts Reported ±1 SE ranges overlap
MedCode 42.47% ±2.16 43.29% ±2.31 0.82 pts Reported ±1 SE ranges overlap
MedScribe 80.17% ±2.00 86.88% ±1.94 6.71 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 50.00% ±5.03 31.00% ±4.65 19.00 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 92.42% ±2.02 92.93% ±1.29 0.50 pts Reported ±1 SE ranges overlap
MMLU Pro 86.97% ±0.34 89.22% ±0.31 2.24 pts Reported ±1 SE ranges do not overlap
Code Migration 41.54% ±4.30 36.60% ±4.14 4.94 pts Reported ±1 SE ranges overlap
IOI 51.61% ±2.40 40.56% ±3.95 11.05 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 87.53% ±0.96 87.35% ±0.96 0.17 pts Reported ±1 SE ranges overlap
ProgramBench 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
SkillsBench 53.83% ±4.49 66.03% ±4.60 12.21 pts Reported ±1 SE ranges do not overlap
SWE-bench 96.40% ±0.83 86.60% ±1.52 9.80 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 14.14% ±2.02 8.59% ±1.34 5.55 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 82.30% ±3.15 69.00% ±4.48 13.30 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 72.86% ±5.50 71.01% ±5.79 1.84 pts Reported ±1 SE ranges overlap

Performance by category

Category DeepSeek V4 Pro 0813 average Grok 4.5 average
Legal 43.58% 45.62%
Finance 58.73% 59.31%
Healthcare 61.32% 65.09%
Math 50.00% 31.00%
Academic 89.70% 91.07%
Coding 53.42% 49.34%
Cyber 72.86% 71.01%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4 Pro 0813 cost Grok 4.5 cost DeepSeek V4 Pro 0813 latency Grok 4.5 latency
Vals Index $3.35 $2.46 1h04m 15m44s
Harvey's Legal Agent Benchmark $0.17 $2.02 21m53s 10m15s
Legal Research Bench $1.09 $1.09 42m07s 15m54s
LegalBench N/A N/A 20.05s 67.88s
EMB $1.08 $1.55 37m27s 18m28s
Finance Agent (v2) $0.88 $1.16 16m15s 4m49s
Tax Agent Bench $0.73 $0.47 26m37s 4m44s
TaxEval v2 N/A N/A 3m46s 3m15s
MedCode N/A N/A 3m01s 59.64s
MedScribe N/A N/A 2m36s 11m45s
ProofBench v1.1 $0.06 $0.57 12m12s 11m09s
GPQA Diamond N/A N/A 4m39s 3m38s
MMLU Pro N/A N/A 59.62s 5m58s
Code Migration $18.58 $3.37 3h32m 23m29s
IOI $2.15 $5.64 1h07m 1h27m
LiveCodeBench N/A N/A 4m33s 9m49s
ProgramBench $0.43 $4.12 1h23m 18m59s
SkillsBench $0.37 $0.58 11m15s 2m16s
SWE-bench $0.10 $0.54 4m00s 3m20s
Terminal-Bench 4.0 $3.31 $5.79 1h32m 33m51s
Vibe Code Bench v1.1 $0.36 $4.10 1h06m 13m50s
CyberBench v1.1 $0.59 $2.95 22m33s 14m08s

Results available only for DeepSeek V4 Pro 0813

  • Terminal-Bench Science
  • Vibe Code Bench 1-100

Results available only for Grok 4.5

  • MortgageTax
  • MMMU Pro
  • SAGE
  • SRE Bench
  • Time Horizon Index: KSP
Model details DeepSeek V4 Pro 0813 Model details Grok 4.5