May 5, 2025
Qwen 3 235B evaluations released!
We just evaluated Qwen 3 235B on all benchmarks!
-
Qwen 3 235B demonstrates exceptional math reasoning capabilities, ranking #3 on Math500, #5 on AIME, and #3 on MGSM.
-
With its βthinking allowedβ approach, Qwen 3 outperforms several prominent closed-source reasoning models including Claude 3.7 Sonnet and o4-mini in mathematical reasoning tasks.
-
Private benchmark challenges: Qwen 3 shows limitations on proprietary benchmarks, particularly struggling on TaxEval where it ranks #29 out of 43 evaluated models.
-
This evaluation showcases Qwen 3βs strong specialized reasoning capabilities while highlighting areas where further improvements could enhance its performance on domain-specific tasks.