Release Date: Oct 15, 2025

Developer Anthropic Β πŸ‡ΊπŸ‡Έ
Context Window 200k
Max Output Tokens 64k
Token Costs (in/out) $1.00/5.00
Weights Private
Input Modalities

Accuracy

22.90 % Β± 0.84

Cost / Test (Vals Index)

$ 0.698

Latency

8 min 20 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.84
52/58

0.0%

Β±1.06
50/61

0.0%

Β±2.39
52/58

0.0%

Β±0.56
53/61

0.0%

Β±2.14
51/61

0.0%

Β±2.00
76/93

0.0%

Β±1.90
23/95

0.0%

Β±0.96
56/98

0.0%

Β±3.13
66/81

0.0%

Β±0.91
108/145

0.0%

Β±3.13
79/96

0.0%

Β±2.47
90/138

0.0%

Β±1.11
130/143

0.0%

Β±0.49
79/145

0.0%

Β±0.48
102/138

0.0%

Β±1.20
91/93

0.0%

Β±0.00
41/47

0.0%

Β±2.11
71/88

0.0%

Β±1.95
57/66
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Anthropic
Temperature: 1
Top P: Default
Top K: Default
Max Output Tokens: 64,000

Updates

Oct 16, 2025

We evaluated Claude Haiku 4.5 (Thinking) and found strong performance:

  • The model places 3rd on our Vals Index, demonstrating well-rounded capabilities across diverse tasks.
  • On Terminal-Bench, Haiku 4.5 achieves 3rd place, showing particular strength on coding tasks.
  • While it performs well on certain coding benchmarks, the model achieves middle-of-the-pack performance on most other benchmarks.
  • The model struggles significantly on our proprietary CaseLaw benchmark and the public MedQA, GPQA Diamond, MMLU Pro, MMMU Pro, and LiveCodeBench benchmarks.
  • Compared to Claude Sonnet 4.5 (Thinking), Haiku trades some performance for significantly faster speed and lower cost - see our model comparison for details.

Overall, Haiku 4.5 sits firmly on the Pareto frontier.