Release Date: Jul 22, 2026

Developer Anthropicย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 128k
Token Costs (in/out) $5.00/25.00
Weights Private
Input Modalities

Accuracy

67.21 % ยฑ 0.98

Cost / Test (Vals Index)

$ 18.81

Latency

55 min 49 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ0.98
2/55

0.0%

2/8

0.0%

ยฑ4.37
2/57

0.0%

ยฑ2.54
22/24

0.0%

ยฑ2.24
3/55

0.0%

ยฑ0.08
7/58

0.0%

ยฑ0.00
1/8

0.0%

ยฑ3.46
2/58

0.0%

ยฑ1.99
1/90

0.0%

ยฑ1.92
2/92

0.0%

ยฑ0.88
1/98

0.0%

ยฑ0.99
3/29

0.0%

ยฑ3.29
19/80

0.0%

ยฑ1.10
1/33

0.0%

ยฑ2.03
3/7

0.0%

ยฑ2.96
2/18

0.0%

ยฑ0.83
23/145

0.0%

ยฑ3.00
4/93

0.0%

ยฑ2.06
2/10

0.0%

ยฑ1.24
8/138

0.0%

ยฑ8.33
1/62

0.0%

ยฑ0.91
4/143

0.0%

ยฑ0.42
7/142

0.0%

ยฑ0.28
2/138

0.0%

ยฑ0.72
2/93

0.0%

ยฑ1.21
3/45

0.0%

ยฑ4.58
7/33

0.0%

ยฑ0.76
1/88

0.0%

ยฑ0.99
4/63
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Compute Effort: max

Updates

Jul 23, 2026

We evaluated Anthropicโ€™s new Claude Opus 5 across 27 benchmark leaderboards.

We ran Opus 5 with Claude Opus 4.8 as a server-side fallback for refusals. Counting fallback-assisted results as failures changes Terminal-Bench 2.1 from 84.64% to 81.27%, MMLU Pro from 91.59% to 91.58%, the Vals Index from 74.82% to 74.47%, and the Vals Multimodal Index from 73.90% to 73.58%. Fallbacks on Finance Agent v2 and CyberBench did not change their published scores.

The model has a 1M-token context window and 128k max output tokens. Evaluations were run with compute effort set to โ€œmaxโ€ except Terminal-Bench 2.1, which used โ€œhighโ€ effort, and temperature set to 1.0 where configurable.

Congrats to the Anthropic team on the release!