May 28, 2026
Anthropic's Claude Opus 4.8 evaluated across our benchmark suite
-
We evaluated Anthropicโs new Claude Opus 4.8 across our benchmark suite. It is the new #1 model on the Vals Index (70.17%) and the Vals Multimodal Index (70.71%), edging out GPT 5.5 on both.
-
Agentic work is the headline. The model is #1 on SWE-bench Verified (88.60%), #1 on Vibe Code Bench (82.72%), and #1 on ProofBench (69.00%). It also places #2 on Terminal Bench 2 (70.04%) and #2 on Finance Agent v2 (53.92%).
-
Knowledge benchmarks land near the top: LCB 87.82% (#3), GPQA 92.42% (#4), MMLU Pro 89.58% (#4), and MMMU 86.59% (#9).
-
Domain-specific benchmarks hold up: MortgageTax 69.91% (#2), MedCode 53.22% (#5), MedScribe 85.75% (#6), TaxEval v2 75.63% (#7), and CorpFin v2 66.71% (#8). Claude Opus 4.8 also picks up #3 on Sage (54.79%).
-
The model still lags behind the field on LegalBench, where it ranks #27 at 83.57%.
The model has a 1M-token context window. Evaluations were run with compute effort set to โmaxโ, a temperature of 1.0, and max output tokens set to 128k.
Congrats to the Anthropic team on the release!