Release Date: Feb 24, 2025

Developer Anthropic ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 200k
Max Output Tokens 8k
Token Costs (in/out) $3.00/15.00
Weights Private
Input Modalities

Accuracy

71.05 %

Avg. Cost (In/Out)

$ 3.00 / $ 15.00

Latency

7.60 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.77
12/94

0.0%

ยฑ0.88
59/139

0.0%

ยฑ2.35
96/132

0.0%

ยฑ1.15
105/136

0.0%

ยฑ0.48
80/136

0.0%

ยฑ0.38
81/132

0.0%

ยฑ1.08
63/88
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 8,192

Updates

Feb 25, 2025

We just evaluated Anthropicโ€™s Claude 3.7 Sonnet (Nonthinking) model!

  • We evaluted the model with Thinking Disabled on all benchmarks. It shows great performance and reaches second place just behind its Thinking Enabled counterpart on Corp Fin.
  • We also evaluated the model with Thinking Enabled. Unlike most models that excel in specific areas, Anthropicโ€™s Claude 3.7 Sonnet (Thinking) demonstrates remarkable consistency, achieving top-tier performance across all evaluated benchmarks. The remaining two benchmarks are currently in progress due to their higher token requirements.

We have also run Google 2.0 Flash Thinking Exp and Google 2.0 Pro Exp on most benchmarks.