Sep 8, 2025
Qwen3 Max Preview Benchmarked
We evaluated Qwen 3 Max Preview on our benchmarks. Despite the modelβs large size, we found the performance did not live up to the hype.
-
On our benchmarks, it was generally in the middle of the pack - but not in the top 5 on any benchmarks, and on most, it was outside the top 20. On Finance Agent, it only managed to get 17% accuracy.
-
Qwen 3 Max Preview did have comparatively strong performance on MGSM and GPQA Diamond, but these benchmarks are saturated, and incremental gains here do not signify meaningful differences in model intelligence.
-
This model is not open source, which is one of the main benefits of the Qwen series. It is also more expensive than its open source counterpart, Qwen 3 (235B), but often performs worse.
Alibaba has currently only released the non-reasoning version of max preview. Weβre excited to benchmark the reasoning version when itβs available, which may improve responses on benchmarks like AIME.