Jul 23, 2026
Anthropic's Claude Opus 5 evaluated across our benchmark suite
We evaluated Anthropicโs new Claude Opus 5 across 27 benchmark leaderboards.
-
Opus 5 places #1 on the Vals Index (67.21%) and #2 on the Vals Multimodal Index (73.90%), within 1.18 and 0.26 points of Claude Fable 5, respectively.
-
Opus 5 takes #1 on 14 leaderboards: SWE-bench Verified (97.00%), IOI (91.67%), Code Migration (57.47%), ProgramBench (3.00% fully resolved), CorpFin v2 (73.19%), Finance Agent v2 (58.63%), Legal Research Bench (55.29%), MedCode (63.57%), MedScribe (90.99%), MortgageTax (72.06%), Public Benefits Bench (76.93%), MMLU Pro (91.59%), MMMU (89.88%), and ProofBench (78.00%).
-
Its other coding results include #2 on Terminal-Bench 2.1 (84.64%), Vibe Code Bench (88.40%), and LiveCodeBench (89.03%).
We ran Opus 5 with Claude Opus 4.8 as a server-side fallback for refusals. Counting fallback-assisted results as failures changes Terminal-Bench 2.1 from 84.64% to 81.27%, MMLU Pro from 91.59% to 91.58%, the Vals Index from 74.82% to 74.47%, and the Vals Multimodal Index from 73.90% to 73.58%. Fallbacks on Finance Agent v2 and CyberBench did not change their published scores.
The model has a 1M-token context window and 128k max output tokens. Evaluations were run with compute effort set to โmaxโ except Terminal-Bench 2.1, which used โhighโ effort, and temperature set to 1.0 where configurable.
Congrats to the Anthropic team on the release!