Dec 3, 2025
Mistral 3 Large evaluated on all benchmarks!
We found that Mistral Large 3 struggles across domains:
- We found a variety of difficulties with agentic use-cases involving tool-calling. Weโre working with the Mistral team to address one such issue.
- The model places second-to-last on our Vals Multimodal Index, ahead of Llama 4 Maverick.
However, there are a few caveats:
- the model represents a meaningful improvement over predecessors like Magistral Medium 1.2 (09/2025) and Mistral Medium 3.1 (05/2025).
- The model is open-weight and performs comparably to other open-weight models like GPT OSS 120B.
We look forward to further improvements from Mistral to come!