Apr 16, 2026
Claude Opus 4.7 is the new SOTA
The newly-released Claude Opus 4.7 has taken first place on both our Vals Index and Vals Multimodal Index, both by substantial margins. Here are our key takeaways:
- Claude Opus 4.7 performs excellently on coding tasks, placing first on Terminal-Bench 2.0 and taking the top spot by several percent on SWE-bench Verified and Vibe Code Bench.
- Claude Opus 4.7 also excels across other domains, placing first on our Finance Agent Benchmark and our math education benchmark, SAGE.
- The model is very consistent, placing in the top ten of all of our benchmarks so far, with the sole exception of Medscribe where it is within 1% of the top ten.
- We also encountered a relatively large number of safety refusals on Claude Opus 4.7, with the model refusing to answer certain questions on benchmarks such as GPQA Diamond and Corpfin.
Congrats to Anthropic on the release!