Model comparison

Claude Sonnet 5 vs DeepSeek V4.1 Flash: Benchmark Comparison

Claude Sonnet 5 has the higher score on 11 of 20 shared benchmarks; DeepSeek V4.1 Flash leads on 9.

The largest observed score gap is 23.33 pts on SkillsBench , where DeepSeek V4.1 Flash leads.

Reported ±1 standard-error ranges overlap on 12 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5 DeepSeek V4.1 Flash Gap Reported uncertainty
Vals Index 51.77% ±1.09 51.32% ±1.13 0.45 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 5.00% ±1.65 6.67% ±1.83 1.67 pts Reported ±1 SE ranges overlap
Legal Research Bench 41.83% ±3.43 41.35% ±3.42 0.48 pts Reported ±1 SE ranges overlap
LegalBench 83.92% ±0.46 83.28% ±0.46 0.64 pts Reported ±1 SE ranges overlap
EMB 66.32% ±3.01 57.21% ±3.23 9.11 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 53.91% ±0.52 53.48% ±0.39 0.43 pts Reported ±1 SE ranges overlap
Tax Agent Bench 62.27% ±3.19 62.46% ±3.15 0.18 pts Reported ±1 SE ranges overlap
MedCode 47.54% ±2.27 41.17% ±2.04 6.37 pts Reported ±1 SE ranges do not overlap
MedScribe 76.05% ±3.05 85.50% ±1.92 9.45 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 77.00% ±4.23 54.00% ±5.01 23.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 2.86% ±2.01 0.00% ±0.00 2.86 pts Reported ±1 SE ranges do not overlap
SAGE 48.92% ±3.40 47.88% ±3.44 1.04 pts Reported ±1 SE ranges overlap
Code Migration 44.39% ±4.25 45.62% ±4.29 1.23 pts Reported ±1 SE ranges overlap
IOI 45.00% ±2.75 40.28% ±2.56 4.72 pts Reported ±1 SE ranges overlap
SkillsBench 46.48% ±4.49 69.80% ±3.88 23.33 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 9.60% ±1.01 19.70% ±1.75 10.10 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 13.82% ±2.98 16.38% ±3.19 2.56 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 81.33% ±3.05 84.74% ±2.89 3.41 pts Reported ±1 SE ranges overlap
CyberBench v1.1 61.91% ±5.74 73.69% ±5.48 11.78 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 66.03% ±1.23 64.28% ±1.25 1.76 pts Reported ±1 SE ranges overlap

Performance by category

Category Claude Sonnet 5 average DeepSeek V4.1 Flash average
Legal 43.58% 43.76%
Finance 60.83% 57.72%
Healthcare 61.80% 63.34%
Math 77.00% 54.00%
Science 2.86% 0.00%
Education 48.92% 47.88%
Coding 40.10% 46.09%
Beta 61.91% 73.69%
Social Mobility 66.03% 64.28%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5 cost DeepSeek V4.1 Flash cost Claude Sonnet 5 latency DeepSeek V4.1 Flash latency
Vals Index $13.72 $0.33 54m16s 27m53s
Harvey's Legal Agent Benchmark $8.95 $0.17 38m52s 9m39s
Legal Research Bench $2.72 $0.25 25m46s 11m42s
LegalBench N/A N/A 4.90s 6.53s
EMB $10.29 $0.23 49m46s 13m00s
Finance Agent (v2) $0.75 $0.21 13m12s 6m04s
Tax Agent Bench $1.79 $0.13 19m19s 5m49s
MedCode N/A N/A 2m15s 35.13s
MedScribe N/A N/A 4m12s 53.83s
ProofBench v1.1 $1.37 $0.13 14m58s 9m07s
Terminal-Bench Science $18.70 $0.47 2h55m 2h57m
SAGE N/A N/A 7m14s 35.78s
Code Migration $35.31 $0.94 1h57m 1h45m
IOI $12.89 $0.24 59m44s 18m30s
SkillsBench $2.90 $0.09 15m12s 3m43s
Terminal-Bench 4.0 $26.33 $0.50 1h45m 48m22s
Vibe Code Bench 1-100 $71.15 $0.69 4h16m 28m21s
Vibe Code Bench v1.1 $25.39 $0.41 1h07m 15m28s
CyberBench v1.1 $1.89 $0.07 16m57s 13m50s
Public Benefits Bench v1.1 $1.29 $0.07 21m09s 9m59s

Results available only for Claude Sonnet 5

  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • ProgramBench
  • SWE-bench

Results available only for DeepSeek V4.1 Flash

  • BioMysteryBench
  • MysteryMechanism
  • SRE Bench
Model details Claude Sonnet 5 Model details DeepSeek V4.1 Flash