Feb 25, 2025
Anthropic's Claude 3.7 Sonnet (Nonthinking) Evaluated on All Benchmarks.
We just evaluated Anthropicโs Claude 3.7 Sonnet (Nonthinking) model!
- We evaluted the model with Thinking Disabled on all benchmarks. It shows great performance and reaches second place just behind its Thinking Enabled counterpart on Corp Fin.
- We also evaluated the model with Thinking Enabled. Unlike most models that excel in specific areas, Anthropicโs Claude 3.7 Sonnet (Thinking) demonstrates remarkable consistency, achieving top-tier performance across all evaluated benchmarks. The remaining two benchmarks are currently in progress due to their higher token requirements.
We have also run Google 2.0 Flash Thinking Exp and Google 2.0 Pro Exp on most benchmarks.