Release Date: Mar 24, 2025

Developer DeepSeek 🇨🇳
Context Window 131k
Max Output Tokens 131k
Token Costs (in/out) $0.90/0.90
Weights Open
Input Modalities

Accuracy

59.51 %

Avg. Cost (In/Out)

$ 0.90 / $ 0.90

Latency

48.11 s

Vals Index
BenchmarksAccuracyRankings

0.0%

±0.88
86/145

0.0%

±2.45
111/138

0.0%

±0.90
58/62

0.0%

±1.16
101/143

0.0%

±0.42
103/142

0.0%

±0.40
96/138
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : DeepSeek
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 131,072

Updates

Mar 26, 2025

We just evaluated DeepSeek V3 on all benchmarks!

  • DeepSeek V3 is DeepSeek’s latest model, boasting speeds of 60 tokens/second and claiming to be 3x faster than V2, with an average accuracy of 73.9% (4.2% better than previous versions).
  • DeepSeek V3 performs comparably (slightly better) to Claude 3.7 Sonnet (71.7%).
  • The model demonstrates strong legal capabilities, scoring particularly well on CaseLaw and LegalBench, though it scores lower on ContractLaw.
  • It shows impressive academic versatility with top-tier performance on MGSM, Math500, and MedQA.