Release Date: Aug 5, 2025

Developer AnthropicΒ πŸ‡ΊπŸ‡Έ
Context Window 200k
Max Output Tokens 32k
Token Costs (in/out) $15.00/75.00
Weights Private
Input Modalities

Accuracy

60.77 %

Avg. Cost (In/Out)

$ 15.00 / $ 75.00

Latency

9 min 39 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±1.96
43/90

0.0%

Β±2.02
77/92

0.0%

Β±0.97
63/98

0.0%

Β±3.02
66/80

0.0%

Β±0.88
79/145

0.0%

Β±2.30
96/138

0.0%

Β±4.91
33/62

0.0%

Β±1.11
102/143

0.0%

Β±0.45
49/142

0.0%

Β±0.33
33/138

0.0%

Β±1.06
58/93
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 32,000

Updates

Aug 8, 2025

We just released results on Claude Opus 4.1 (Nonthinking) and found that, despite achieving top spots on MMLU Pro and MGSM, the model performs only marginally better across almost all of our benchmarks (<2% performance gain) compared to Claude Opus 4 (Nonthinking).

On our private benchmarks, Opus 4.1 fails to place among the top 10 models. On public benchmarks, however, the model breaks the top 10 on 5 of the 9 public benchmarks we evaluated. This signals the need for more private benchmarks to evaluate meaningful differences between models and gauge true performance.