Release Date: Apr 28, 2025

Developer Alibabaย ๐Ÿ‡จ๐Ÿ‡ณ
Context Window 128k
Max Output Tokens 33k
Token Costs (in/out) $0.22/0.88
Weights Open
Input Modalities

Accuracy

74.58 %

Avg. Cost (In/Out)

$ 0.22 / $ 0.88

Latency

2 min 59 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ0.89
92/145

0.0%

ยฑ2.37
95/138

0.0%

ยฑ1.15
87/143

0.0%

ยฑ0.44
86/144

0.0%

ยฑ0.39
84/138
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Alibaba
Temperature: 0.7
Top P: Default
Top K: Default
Max Output Tokens: 32,768

Updates

May 5, 2025

We just evaluated Qwen 3 235B on all benchmarks!

  • Qwen 3 235B demonstrates exceptional math reasoning capabilities, ranking #3 on Math500, #5 on AIME, and #3 on MGSM.

  • With its โ€œthinking allowedโ€ approach, Qwen 3 outperforms several prominent closed-source reasoning models including Claude 3.7 Sonnet and o4-mini in mathematical reasoning tasks.

  • Private benchmark challenges: Qwen 3 shows limitations on proprietary benchmarks, particularly struggling on TaxEval where it ranks #29 out of 43 evaluated models.

  • This evaluation showcases Qwen 3โ€™s strong specialized reasoning capabilities while highlighting areas where further improvements could enhance its performance on domain-specific tasks.