Release Date: Jul 22, 2026

Developer Anthropicย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 128k
Token Costs (in/out) $5.00/25.00
Weights Private
Input Modalities

Accuracy

67.21 % ยฑ 0.98

Cost / Test (Vals Index)

$ 19.28

Latency

55 min 49 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ0.98
1/46

0.0%

ยฑ0.00
1/3

0.0%

ยฑ4.37
1/49

0.0%

ยฑ2.54
20/22

0.0%

ยฑ2.24
2/46

0.0%

ยฑ0.08
3/49

0.0%

ยฑ0.00
1/7

0.0%

ยฑ3.46
1/49

0.0%

ยฑ1.99
1/85

0.0%

ยฑ1.92
1/84

0.0%

ยฑ0.88
1/95

0.0%

ยฑ0.99
1/23

0.0%

ยฑ3.29
18/76

0.0%

ยฑ1.10
1/30

0.0%

ยฑ2.03
2/5

0.0%

ยฑ0.83
21/140

0.0%

ยฑ3.00
2/84

0.0%

ยฑ1.24
7/133

0.0%

ยฑ8.33
1/62

0.0%

ยฑ0.91
2/138

0.0%

ยฑ0.42
5/137

0.0%

ยฑ0.28
1/133

0.0%

ยฑ0.72
1/89

0.0%

ยฑ1.21
1/39

0.0%

ยฑ4.58
7/28

0.0%

ยฑ0.76
1/83

0.0%

ยฑ0.99
2/54
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Compute Effort: max

Updates

Jul 23, 2026

We evaluated Anthropicโ€™s new Claude Opus 5 across 27 benchmark leaderboards.

We ran Opus 5 with Claude Opus 4.8 as a server-side fallback for refusals. Counting fallback-assisted results as failures changes Terminal-Bench 2.1 from 84.64% to 81.27%, MMLU Pro from 91.59% to 91.58%, the Vals Index from 74.82% to 74.47%, and the Vals Multimodal Index from 73.90% to 73.58%. Fallbacks on Finance Agent v2 and CyberBench did not change their published scores.

The model has a 1M-token context window and 128k max output tokens. Evaluations were run with compute effort set to โ€œmaxโ€ except Terminal-Bench 2.1, which used โ€œhighโ€ effort, and temperature set to 1.0 where configurable.

Congrats to the Anthropic team on the release!