Release Date: Apr 28, 2025

Developer Alibaba πŸ‡¨πŸ‡³
Context Window 128k
Max Output Tokens 33k
Token Costs (in/out) $0.22/0.88
Weights Open
Input Modalities

Accuracy

62.15 %

Avg. Cost (In/Out)

$ 0.22 / $ 0.88

Latency

3 min 28 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.89
86/139

0.0%

Β±2.37
89/132

0.0%

Β±0.00
62/62

0.0%

Β±1.15
80/136

0.0%

Β±0.44
78/136

0.0%

Β±0.39
78/132
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Alibaba
Temperature: 0.7
Top P: Default
Top K: Default
Max Output Tokens: 32,768

Updates

May 5, 2025

We just evaluated Qwen 3 235B on all benchmarks!

  • Qwen 3 235B demonstrates exceptional math reasoning capabilities, ranking #3 on Math500, #5 on AIME, and #3 on MGSM.

  • With its β€œthinking allowed” approach, Qwen 3 outperforms several prominent closed-source reasoning models including Claude 3.7 Sonnet and o4-mini in mathematical reasoning tasks.

  • Private benchmark challenges: Qwen 3 shows limitations on proprietary benchmarks, particularly struggling on TaxEval where it ranks #29 out of 43 evaluated models.

  • This evaluation showcases Qwen 3’s strong specialized reasoning capabilities while highlighting areas where further improvements could enhance its performance on domain-specific tasks.