Release Date: May 22, 2025

Developer AnthropicΒ πŸ‡ΊπŸ‡Έ
Context Window 200k
Max Output Tokens 32k
Token Costs (in/out) $15.00/75.00
Weights Private
Input Modalities

Accuracy

72.48 %

Avg. Cost (In/Out)

$ 15.00 / $ 75.00

Latency

13.01 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.98
70/98

0.0%

Β±0.88
72/145

0.0%

Β±2.26
92/138

0.0%

Β±1.12
104/143

0.0%

Β±0.44
55/144

0.0%

Β±0.34
48/138

0.0%

Β±1.06
60/93
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 32,000

Updates

May 30, 2025

We’ve released our evaluation of Claude Opus 4 (Nonthinking) across our benchmarks!

We found:

  • Opus 4 ranks #1 on both MMLU Pro and MGSM, narrowly setting new state-of-the-art scores. However, it achieves middle of the road performance across most other benchmarks.
  • Compared to its predecessor (Opus 3), Opus 4 ranked higher on CaseLaw (#22 vs. #24/62) and LegalBench (#8 vs #32/67) but scored notably lower on ContractLaw (#16 vs. #2/69)
  • Opus 4 is expensive, with an output cost of $75.00 /M tokens, 5x as much as Sonnet 4, and about 1.5x more expensive than o3 ($15 / $75 vs $10 / $40).

We also benchmarked Claude Sonnet 4 (Thinking) and Claude Sonnet 4 (Nonthinking) on our Finance Agent benchmark (the last remaining benchmark for this model). They performed nearly identically to Claude 3.7 Sonnet (Nonthinking).