May 25, 2025
Claude Sonnet 4 (Nonthinking) evaluated on all benchmarks!
We just evaluated Claude Sonnet 4 (Nonthinking) on all benchmarks!
- Claude Sonnet 4 (Nonthinking) achieves 76.9% accuracy on average, a 7.1% improvement on Anthropicโs previous flagship model, Claude 3.7 Sonnet. The newer Claude is also nearly twice as fast for the same price.
- Claude Sonnet 4 (Nonthinking) excels on the MGSM benchmark, edging out Claude 3.7 Sonnet (Thinking) by a tenth of a percentage point.
- Claude Sonnet 4 (Nonthinking) also achieves strong performance on our proprietary CaseLaw benchmark, outperforming all previous Anthropic models.
- Interestingly, Claude Sonnet 4 (Nonthinking) performs worse than its predecessor Claude 3.7 Sonnet by six percentage points on the MortgageTax benchmark. It even performs worse than its predecessor, Claude 3.5 Sonnet, on both the MortgageTax and CorpFin benchmarks!
Stay tuned for evaluations of Sonnet 4โs thinking variant, as well as Opus 4!