Feb 17, 2026
Claude Sonnet 4.6 - Anthropic's latest model
Anthropicโs latest model, Claude Sonnet 4.6, debuts at #1 on our Vals Index and #1 on the Vals Multimodal Index, evaluated across 17 benchmarks. Key takeaways:
- Claude Sonnet 4.6 takes first place on both Finance Agent (63.3%) and Tax Eval v2 (77.1%), demonstrating strong domain expertise in finance and tax.
- It takes first place on Terminal-Bench 2.0 (59.55%), beating out Claude Opus 4.5 (Thinking) and Claude Opus 4.6 (Thinking). It is also #3 on SWE-bench Verified (76.2%) and #2 on ProofBench (45.0%), showing competitive performance on agentic coding and formal mathematics.
- On knowledge benchmarks, Sonnet 4.6 scores 92.3% on AIME, 87.3% on MMLU Pro (#5), and 85.6% on GPQA Diamond.
- The model finishes top 10 on 13 of 17 benchmarks, with particularly strong results across finance and legal domains including #6 on CaseLaw and #7 on Mortgage Tax.