DeepSeek V4 vs Kimi K2.6: Benchmark Comparison

DeepSeek V4 has the higher score on 9 of 17 shared benchmarks; Kimi K2.6 leads on 7.

The largest observed score gap is 13.36 pts on Tax Agent Bench , where DeepSeek V4 leads.

Reported ±1 standard-error ranges overlap on 10 of 17 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4 Kimi K2.6 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 3.75% ±1.43 1.67% ±0.83 2.08 pts Reported ±1 SE ranges overlap
Legal Research Bench 23.08% ±2.93 15.87% ±2.54 7.21 pts Reported ±1 SE ranges do not overlap
LegalBench 80.32% ±0.47 84.74% ±0.45 4.42 pts Reported ±1 SE ranges do not overlap
EMB 51.62% ±3.05 57.85% ±2.74 6.23 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 44.08% ±0.65 44.90% ±0.73 0.82 pts Reported ±1 SE ranges overlap
Tax Agent Bench 58.50% ±3.13 45.13% ±3.23 13.36 pts Reported ±1 SE ranges do not overlap
TaxEval v2 72.08% ±0.88 74.65% ±0.85 2.58 pts Reported ±1 SE ranges do not overlap
MedCode 40.45% ±2.12 40.14% ±2.04 0.31 pts Reported ±1 SE ranges overlap
MedScribe 75.14% ±2.00 78.15% ±1.79 3.00 pts Reported ±1 SE ranges overlap
GPQA Diamond 89.39% ±1.65 89.14% ±2.00 0.25 pts Reported ±1 SE ranges overlap
MMLU Pro 87.25% ±0.34 87.57% ±0.33 0.32 pts Reported ±1 SE ranges overlap
Code Migration 26.20% ±4.04 27.77% ±4.14 1.57 pts Reported ±1 SE ranges overlap
LiveCodeBench 87.48% ±0.95 86.77% ±0.97 0.71 pts Reported ±1 SE ranges overlap
ProgramBench 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
SWE-bench 77.40% ±1.87 76.20% ±1.91 1.20 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 49.93% ±4.77 37.89% ±4.91 12.04 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 62.92% ±1.26 56.63% ±1.29 6.29 pts Reported ±1 SE ranges do not overlap

Performance by category

Category DeepSeek V4 average Kimi K2.6 average
Legal 35.72% 34.09%
Finance 56.57% 55.64%
Healthcare 57.80% 59.15%
Academic 88.32% 88.36%
Coding 48.20% 45.73%
Social Mobility 62.92% 56.63%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4 cost Kimi K2.6 cost DeepSeek V4 latency Kimi K2.6 latency
Harvey's Legal Agent Benchmark $0.74 $0.82 6m42s 23m23s
Legal Research Bench $0.68 $1.52 17m05s 30m53s
LegalBench N/A N/A 3m04s 50.07s
EMB $0.80 $2.20 13m42s 39m48s
Finance Agent (v2) $0.88 $1.14 25m26s 27m10s
Tax Agent Bench $0.72 $0.57 20m45s 15m24s
TaxEval v2 N/A N/A 3m10s 4m27s
MedCode N/A N/A 6m23s 5m05s
MedScribe N/A N/A 5m46s 8m14s
GPQA Diamond N/A N/A 15m00s 9m05s
MMLU Pro N/A N/A 9m38s 2m36s
Code Migration $1.07 $5.11 28m58s 1h30m
LiveCodeBench N/A N/A 11m56s 8m49s
ProgramBench N/A N/A 1h03m 1h20m
SWE-bench $0.44 $0.27 10m35s 8m40s
Vibe Code Bench v1.1 $2.21 $1.93 56m59s 49m28s
Public Benefits Bench v1.1 $0.46 $0.91 15m22s 9m55s

Results available only for DeepSeek V4

  • Vals Index
  • ProofBench v1.1
  • SkillsBench
  • Terminal-Bench 4.0

Results available only for Kimi K2.6

  • MortgageTax
  • MMMU Pro
  • SAGE
Model details DeepSeek V4 Model details Kimi K2.6