Release Date: Jun 30, 2026

Developer Anthropicย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 1M
Max Output Tokens 128k
Token Costs (in/out) $3.00/15.00
Weights Private
Input Modalities

Accuracy

59.61 % ยฑ 1.15

Cost / Test (Vals Index)

$ 17.70

Latency

45 min 46 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ1.15
6/46

0.0%

ยฑ4.25
8/48

0.0%

ยฑ3.01
9/46

0.0%

ยฑ0.52
11/49

0.0%

ยฑ3.43
9/49

0.0%

ยฑ2.27
22/84

0.0%

ยฑ3.05
57/83

0.0%

ยฑ0.90
3/94

0.0%

ยฑ4.21
5/23

0.0%

ยฑ3.40
20/75

0.0%

ยฑ1.23
9/30

0.0%

ยฑ0.84
14/139

0.0%

ยฑ3.05
6/84

0.0%

ยฑ2.22
29/132

0.0%

ยฑ1.09
49/137

0.0%

ยฑ0.46
37/136

0.0%

ยฑ0.37
23/132

0.0%

ยฑ0.90
30/88

0.0%

ยฑ0.00
13/38

0.0%

ยฑ4.49
21/27

0.0%

ยฑ1.80
21/82

0.0%

ยฑ2.08
9/54
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Compute Effort: max

Updates

Jun 30, 2026

  • We evaluated Anthropicโ€™s new Claude Sonnet 5 across our benchmark suite. It comes in at #3 on the Vals Index (68.61%), behind only Claude Fable 5 (75.15%) and Claude Opus 4.8 (70.36%), and narrowly ahead of GPT 5.5 (67.95%).

  • Itโ€™s a substantial generational step: +8.5 points on the Vals Index over Claude Sonnet 4.6 (60.07%) โ€” and almost all of that gain comes from coding.

  • The models main strength is in coding. Compared with Sonnet 4.6, Claude Sonnet 5 jumps +30.7 points on Vibe Code Bench (56.22% โ†’ 86.90%) and +17.2 points on Terminal-Bench 2.1 (57.30% โ†’ 74.53%), while SWE-bench Verified actually slips slightly (77.45% โ†’ 75.49%). The gains are concentrated in sustained, multi-step agentic coding rather than one-shot patch generation.

  • Outside coding itโ€™s a quieter step. Sonnet 5 posts 67.95% on CorpFin v2 (+1.3 vs Sonnet 4.6) and 51.98% on Finance Agent (+0.9). Its index score clears last-generation Claude Opus 4.7 (66.10%).

We noticed a small but noticeable number of refusals on CorpFin v2 โ€” 15 of 858 tasks, mostly flagged as โ€œbio.โ€

The model has a 1M-token context window. Evaluations were run with compute effort set to โ€œmaxโ€ on all benchmarks except Terminal-Bench 2.1 (run at โ€œhighโ€), 128k max output tokens, and default temperature and top-p. Terminal-Bench 2.1 was averaged over 3 runs (71.9%โ€“78.7%).

Congrats to the Anthropic team on the release!