Release Date: Aug 13, 2026

Developer DeepSeek πŸ‡¨πŸ‡³
Context Window 1M
Max Output Tokens 384k
Token Costs (in/out) $0.43/0.87
Weights Open
Input Modalities

Accuracy

52.37 % Β± 1.14

Cost / Test (Vals Index)

$ 0.843

Latency

58 min 18 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±1.14
18/43

0.0%

Β±4.30
10/48

0.0%

Β±3.06
25/45

0.0%

Β±0.25
18/48

0.0%

Β±3.42
11/46

0.0%

Β±2.16
34/84

0.0%

Β±2.00
38/83

0.0%

Β±0.87
51/139

0.0%

Β±3.15
5/83

0.0%

Β±2.02
14/132

0.0%

Β±0.44
60/136

0.0%

Β±0.34
33/132

0.0%

Β±0.83
2/82

0.0%

Β±1.50
33/53
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : DeepSeek
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 384,000
Reasoning Effort: max

Updates

Aug 12, 2026

We evaluated DeepSeek’s DeepSeek V4 Pro 0813 across our benchmark suite.

  • It ranks #12 on the Vals Index (66.25%), up 10.63 points from DeepSeek V4 (55.62%).

  • Its standout result is SWE-bench Verified, where it places #2 of 82 models (96.40%) and is the highest-scoring open-weight model on the board, ahead of Kimi K3 (93.40%). It is also by far the cheapest model near the top: $0.02 per test, versus $1.29 for Claude Opus 5 at 97.00%.

  • Reasoning and legal results improve sharply over the previous release: 49.00% on ProofBench (up from 10.00%, #45 β†’ #15) and 40.87% on Legal Research Bench (up from 23.08%, #27 β†’ #11). It reaches 7.50% on Harvey’s Legal Agent Benchmark with an 88.06% criteria pass rate, #10 of 43.

  • It struggles on terminal-driven coding and Excel tasks: it scores 54.68% across three full trials of Terminal-Bench 2.1 (#33 of 52, 28.89% on hard tasks) and 52.80% on EMB (#24 of 37).

The model has a 1M-token context window and supports up to 384k output tokens and tool calling. Evaluations were run at max reasoning effort; the model does not accept a temperature parameter.

Congrats to the DeepSeek team on the release!