Qwen 3 Max Preview

Release Date: Sep 5, 2025

Developer AlibabaΒ πŸ‡¨πŸ‡³
Context Window 262k
Max Output Tokens 66k
Token Costs (in/out) $1.20/6.00
Weights Private
Input Modalities

Accuracy

65.04 %

Avg. Cost (In/Out)

$ 1.20 / $ 6.00

Latency

10 min 7 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.86
40/141

0.0%

Β±2.09
72/133

0.0%

Β±3.17
37/62

0.0%

Β±1.17
93/138

0.0%

Β±0.40
76/137

0.0%

Β±0.36
68/133
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Alibaba
Temperature: 0.7
Top P: Default
Top K: Default
Max Output Tokens: 65,536

Updates

Sep 8, 2025

We evaluated Qwen 3 Max Preview on our benchmarks. Despite the model’s large size, we found the performance did not live up to the hype.

  • On our benchmarks, it was generally in the middle of the pack - but not in the top 5 on any benchmarks, and on most, it was outside the top 20. On Finance Agent, it only managed to get 17% accuracy.

  • Qwen 3 Max Preview did have comparatively strong performance on MGSM and GPQA Diamond, but these benchmarks are saturated, and incremental gains here do not signify meaningful differences in model intelligence.

  • This model is not open source, which is one of the main benefits of the Qwen series. It is also more expensive than its open source counterpart, Qwen 3 (235B), but often performs worse.

Alibaba has currently only released the non-reasoning version of max preview. We’re excited to benchmark the reasoning version when it’s available, which may improve responses on benchmarks like AIME.