Release Date: Aug 13, 2026

Developer DeepSeekย ๐Ÿ‡จ๐Ÿ‡ณ
Context Window 1M
Max Output Tokens 384k
Token Costs (in/out) $0.43/0.87
Weights Open
Input Modalities

Accuracy

52.37 % ยฑ 1.14

Cost / Test (Vals Index)

$ 0.843

Latency

58 min 18 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.14
19/46

0.0%

ยฑ4.30
10/49

0.0%

ยฑ3.06
26/46

0.0%

ยฑ0.25
19/49

0.0%

ยฑ3.42
11/49

0.0%

ยฑ2.16
35/85

0.0%

ยฑ2.00
39/84

0.0%

ยฑ5.03
12/23

0.0%

ยฑ0.87
52/140

0.0%

ยฑ3.15
5/84

0.0%

ยฑ2.02
15/133

0.0%

ยฑ0.96
11/138

0.0%

ยฑ0.44
61/137

0.0%

ยฑ0.34
34/133

0.0%

ยฑ0.00
12/39

0.0%

ยฑ4.49
13/28

0.0%

ยฑ0.83
2/83

0.0%

ยฑ1.50
34/54
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : DeepSeek
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 384,000
Reasoning Effort: max

Updates

Aug 12, 2026

We evaluated DeepSeekโ€™s DeepSeek V4 Pro 0813 across our benchmark suite.

  • It ranks #18 on the Vals Index (52.37%), up 9.48 points from DeepSeek V4 (42.89%).

  • Its standout result is SWE-bench Verified, where it places #2 of 82 models (96.40%) and is the highest-scoring open-weight model on the board, ahead of Kimi K3 (93.40%). It is also by far the cheapest model near the top: $0.02 per test, versus $1.29 for Claude Opus 5 at 97.00%.

  • Reasoning and legal results improve sharply over the previous release: 49.00% on ProofBench (up from 10.00%, #45 โ†’ #15) and 40.87% on Legal Research Bench (up from 23.08%, #27 โ†’ #11). It reaches 7.50% on Harveyโ€™s Legal Agent Benchmark with an 88.06% criteria pass rate, #10 of 43.

  • It struggles on terminal-driven coding and Excel tasks: it scores 54.68% across three full trials of Terminal-Bench 2.1 (#33 of 52, 28.89% on hard tasks) and 52.80% on EMB (#24 of 37).

The model has a 1M-token context window and supports up to 384k output tokens and tool calling. Evaluations were run at max reasoning effort; the model does not accept a temperature parameter.

Congrats to the DeepSeek team on the release!