DeepSeek V4.1 Flash vs Kimi K3: Benchmark Comparison

DeepSeek V4.1 Flash has the higher score on 4 of 21 shared benchmarks; Kimi K3 leads on 17.

The largest observed score gap is 33.00 pts on ProofBench v1.1 , where Kimi K3 leads.

Reported ±1 standard-error ranges overlap on 11 of 21 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark DeepSeek V4.1 Flash Kimi K3 Gap Reported uncertainty
Vals Index 51.32% ±1.13 50.30% ±0.99 1.02 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 6.67% ±1.83 12.92% ±2.68 6.25 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.35% ±3.42 46.15% ±3.46 4.81 pts Reported ±1 SE ranges overlap
LegalBench 83.28% ±0.46 86.21% ±0.41 2.94 pts Reported ±1 SE ranges do not overlap
EMB 57.21% ±3.23 66.68% ±2.87 9.47 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 53.48% ±0.39 53.11% ±0.37 0.37 pts Reported ±1 SE ranges overlap
Tax Agent Bench 62.46% ±3.15 68.67% ±3.08 6.21 pts Reported ±1 SE ranges overlap
MedCode 41.17% ±2.04 49.36% ±2.20 8.18 pts Reported ±1 SE ranges do not overlap
MedScribe 85.50% ±1.92 88.05% ±1.98 2.55 pts Reported ±1 SE ranges overlap
ProofBench v1.1 54.00% ±5.01 87.00% ±3.38 33.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 67.78% ±1.11 72.59% ±1.61 4.81 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 0.00% ±0.00 1.43% ±1.43 1.43 pts Reported ±1 SE ranges overlap
SAGE 47.88% ±3.44 52.78% ±3.41 4.90 pts Reported ±1 SE ranges overlap
Code Migration 45.62% ±4.29 16.10% ±4.10 29.52 pts Reported ±1 SE ranges do not overlap
IOI 40.28% ±2.56 48.94% ±9.82 8.67 pts Reported ±1 SE ranges overlap
ProgramBench 0.50% ±0.50 2.00% ±0.99 1.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 19.70% ±1.75 17.17% ±0.51 2.52 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 16.38% ±3.19 18.24% ±3.92 1.87 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 84.74% ±2.89 84.97% ±2.75 0.23 pts Reported ±1 SE ranges overlap
CyberBench v1.1 73.69% ±5.48 75.24% ±5.56 1.55 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 64.28% ±1.25 68.20% ±1.21 3.92 pts Reported ±1 SE ranges do not overlap

Performance by category

Category DeepSeek V4.1 Flash average Kimi K3 average
Legal 43.76% 48.43%
Finance 57.72% 62.82%
Healthcare 63.34% 68.70%
Math 54.00% 87.00%
Science 33.89% 37.01%
Education 47.88% 52.78%
Coding 34.54% 31.24%
Cyber 73.69% 75.24%
Social Mobility 64.28% 68.20%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark DeepSeek V4.1 Flash cost Kimi K3 cost DeepSeek V4.1 Flash latency Kimi K3 latency
Vals Index $0.33 $6.38 27m53s 1h08m
Harvey's Legal Agent Benchmark $0.17 $3.83 9m39s 13m58s
Legal Research Bench $0.25 $3.47 11m42s 12m40s
LegalBench N/A N/A 6.53s 7.11s
EMB $0.23 $3.07 13m00s 12m21s
Finance Agent (v2) $0.21 $1.91 6m04s 4m46s
Tax Agent Bench $0.13 $2.79 5m49s 40m15s
MedCode N/A N/A 35.13s 36.99s
MedScribe N/A N/A 53.83s 38.24s
ProofBench v1.1 $0.13 $1.65 9m07s 34m31s
BioMysteryBench $0.23 $1.06 14m50s 4m55s
Terminal-Bench Science $0.47 $21.30 2h57m 4h49m
SAGE N/A N/A 35.78s 47.30s
Code Migration $0.94 $13.87 1h45m 4h18m
IOI $0.24 $14.17 18m30s 4h18m
ProgramBench $0.90 $70.48 1h02m 5h42m
Terminal-Bench 4.0 $0.50 $12.02 48m22s 3h09m
Vibe Code Bench 1-100 $0.69 $9.40 28m21s 1h48m
Vibe Code Bench v1.1 $0.41 $10.01 15m28s 16m39s
CyberBench v1.1 $0.07 $2.13 13m50s 30m44s
Public Benefits Bench v1.1 $0.07 $0.91 9m59s 7m19s

Results available only for DeepSeek V4.1 Flash

  • MysteryMechanism
  • SkillsBench
  • SRE Bench

Results available only for Kimi K3

  • Vals RSI Index
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SWE-bench
  • Time Horizon Index: KSP
Model details DeepSeek V4.1 Flash Model details Kimi K3