Qwen 3 Max Thinking

Release Date: Jan 23, 2026

Developer Alibaba πŸ‡¨πŸ‡³
Context Window 256k
Max Output Tokens 32k
Token Costs (in/out) $1.20/6.00
Weights Private
Input Modalities

Accuracy

51.31 %

Avg. Cost (In/Out)

$ 1.20 / $ 6.00

Latency

13 min 21 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±1.89
75/84

0.0%

Β±1.91
65/83

0.0%

Β±1.80
45/132

0.0%

Β±3.55
32/62

0.0%

Β±0.35
55/132
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Alibaba
Temperature: 0.7
Top P: Default
Top K: Default
Max Output Tokens: 32,000

Updates

Feb 4, 2026

We found that Qwen 3 Max Thinking shines on financial tasks, but struggles on coding evaluations when compared to its predecessor, Qwen 3 Max. We attribute the difference to use of new β€œreasoning” mode like those featured by closed-source providers.

Its most impressive finish is second on CorpFin, behind Kimi K2.5. It also improves by 10% over its predecessor on our Finance Agent Benchmark.

By contrast, it performs worse than Qwen 3 Max on both SWE-bench Verified and Terminal-Bench 2.0, placing 20th and 25th, respectively.

We also found that the model can get expensive on long-context agentic tasks, with the highest pricing tier equivalent to that of the Claude Sonnet models. Additionally, the model is worse than its predecessor at context caching, which increases prices per token further.

Congrats to the team at Alibaba on the release!