Release Date: May 22, 2025

Developer AnthropicΒ πŸ‡ΊπŸ‡Έ
Context Window 200k
Max Output Tokens 32k
Token Costs (in/out) $15.00/75.00
Weights Private
Input Modalities

Accuracy

72.48 %

Avg. Cost (In/Out)

$ 15.00 / $ 75.00

Latency

13.01 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.98
67/95

0.0%

Β±0.88
68/141

0.0%

Β±2.26
87/133

0.0%

Β±1.12
99/138

0.0%

Β±0.44
50/137

0.0%

Β±0.34
45/133

0.0%

Β±1.06
56/89
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 32,000

Updates

May 30, 2025

We’ve released our evaluation of Claude Opus 4 (Nonthinking) across our benchmarks!

We found:

  • Opus 4 ranks #1 on both MMLU Pro and MGSM, narrowly setting new state-of-the-art scores. However, it achieves middle of the road performance across most other benchmarks.
  • Compared to its predecessor (Opus 3), Opus 4 ranked higher on CaseLaw (#22 vs. #24/62) and LegalBench (#8 vs #32/67) but scored notably lower on ContractLaw (#16 vs. #2/69)
  • Opus 4 is expensive, with an output cost of 75.00/Mtokens,5xasmuchasSonnet4,andabout1.5xmoreexpensivethano3(75.00 /M tokens, 5x as much as Sonnet 4, and about 1.5x more expensive than o3 (15 / 75vs75 vs 10 / $40).

We also benchmarked Claude Sonnet 4 (Thinking) and Claude Sonnet 4 (Nonthinking) on our Finance Agent benchmark (the last remaining benchmark for this model). They performed nearly identically to Claude 3.7 Sonnet (Nonthinking).