Release Date: Jun 30, 2026

Developer Anthropic Β πŸ‡ΊπŸ‡Έ
Context Window 1M
Max Output Tokens 128k
Token Costs (in/out) $2.00/10.00
Weights Private
Input Modalities

Accuracy

51.77 % Β± 1.09

Cost / Test (Vals Index)

$ 13.72

Latency

54 min 16 s

Vals Index
BenchmarksAccuracyRankings

51.77%

Β±1.09
18/30

44.39%

Β±4.25
19/70

61.91%

Β±5.74
30/42

66.32%

Β±3.01
18/68

53.91%

Β±0.52
24/72

41.83%

Β±3.43
19/71

47.54%

Β±2.27
30/103

76.05%

Β±3.05
74/105

70.03%

Β±0.90
4/98

77.00%

Β±4.23
11/43

48.92%

Β±3.40
23/89

66.03%

Β±1.23
16/44

62.27%

Β±3.19
27/63

75.63%

Β±0.84
15/145

13.82%

Β±2.98
16/20

81.33%

Β±3.05
18/105

88.89%

Β±2.22
32/138

45.00%

Β±2.75
26/37

82.43%

Β±1.09
53/143

83.92%

Β±0.46
43/148

87.55%

Β±0.37
26/138

83.01%

Β±0.90
35/93

0.00%

Β±0.00
24/53

46.48%

Β±4.49
26/35

79.60%

Β±1.80
27/88

9.60%

Β±1.01
26/31

2.86%

Β±2.01
22/33
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 128,000
Compute Effort: max

Compared with

Updates

Jun 30, 2026

  • We evaluated Anthropic’s new Claude Sonnet 5 across our benchmark suite. It comes in at #3 on the Vals Index (68.61%), behind only Claude Fable 5 (75.15%) and Claude Opus 4.8 (70.36%), and narrowly ahead of GPT 5.5 (67.95%).

  • It’s a substantial generational step: +8.5 points on the Vals Index over Claude Sonnet 4.6 (60.07%) β€” and almost all of that gain comes from coding.

  • The models main strength is in coding. Compared with Sonnet 4.6, Claude Sonnet 5 jumps +30.7 points on Vibe Code Bench (56.22% β†’ 86.90%) and +17.2 points on Terminal-Bench 2.1 (57.30% β†’ 74.53%), while SWE-bench Verified actually slips slightly (77.45% β†’ 75.49%). The gains are concentrated in sustained, multi-step agentic coding rather than one-shot patch generation.

  • Outside coding it’s a quieter step. Sonnet 5 posts 67.95% on CorpFin v2 (+1.3 vs Sonnet 4.6) and 51.98% on Finance Agent (+0.9). Its index score clears last-generation Claude Opus 4.7 (66.10%).

We noticed a small but noticeable number of refusals on CorpFin v2 β€” 15 of 858 tasks, mostly flagged as β€œbio.”

The model has a 1M-token context window. Evaluations were run with compute effort set to β€œmax” on all benchmarks except Terminal-Bench 2.1 (run at β€œhigh”), 128k max output tokens, and default temperature and top-p. Terminal-Bench 2.1 was averaged over 3 runs (71.9%–78.7%).

Congrats to the Anthropic team on the release!