Aug 12, 2026
DeepSeek's V4 Pro 0813 evaluated across our benchmark suite
We evaluated DeepSeekβs DeepSeek V4 Pro 0813 across our benchmark suite.
-
It ranks #12 on the Vals Index (66.25%), up 10.63 points from DeepSeek V4 (55.62%).
-
Its standout result is SWE-bench Verified, where it places #2 of 82 models (96.40%) and is the highest-scoring open-weight model on the board, ahead of Kimi K3 (93.40%). It is also by far the cheapest model near the top: $0.02 per test, versus $1.29 for Claude Opus 5 at 97.00%.
-
Reasoning and legal results improve sharply over the previous release: 49.00% on ProofBench (up from 10.00%, #45 β #15) and 40.87% on Legal Research Bench (up from 23.08%, #27 β #11). It reaches 7.50% on Harveyβs Legal Agent Benchmark with an 88.06% criteria pass rate, #10 of 43.
-
It struggles on terminal-driven coding and Excel tasks: it scores 54.68% across three full trials of Terminal-Bench 2.1 (#33 of 52, 28.89% on hard tasks) and 52.80% on EMB (#24 of 37).
The model has a 1M-token context window and supports up to 384k output tokens and tool calling. Evaluations were run at max reasoning effort; the model does not accept a temperature parameter.
Congrats to the DeepSeek team on the release!