Release Date: Dec 2, 2025

Developer Mistralย ๐Ÿ‡ซ๐Ÿ‡ท
Context Window 256k
Max Output Tokens 256k
Token Costs (in/out) $0.50/1.50
Weights Open
Input Modalities

Accuracy

50.28 %

Avg. Cost (In/Out)

$ 0.50 / $ 1.50

Latency

12 min 37 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ0.99
83/98

0.0%

ยฑ2.81
77/80

0.0%

ยฑ0.87
54/145

0.0%

ยฑ2.37
100/138

0.0%

ยฑ1.46
50/62

0.0%

ยฑ1.15
113/143

0.0%

ยฑ0.46
93/142

0.0%

ยฑ0.47
94/138

0.0%

ยฑ1.14
74/93

0.0%

ยฑ2.21
86/88
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Mistral
Temperature: 0.065
Top P: Default
Top K: Default
Max Output Tokens: 256,000

Updates

Dec 3, 2025

We found that Mistral Large 3 struggles across domains:

  • We found a variety of difficulties with agentic use-cases involving tool-calling. Weโ€™re working with the Mistral team to address one such issue.
  • The model places second-to-last on our Vals Multimodal Index, ahead of Llama 4 Maverick.

However, there are a few caveats:

We look forward to further improvements from Mistral to come!

Dec 2, 2025

Mistral Large 3 is #20 on our Vals Index (out of 24).

A main source of failure was the modelโ€™s difficulty in calling tools - a common error mode was the model making a function call with a malformed tool name (e.g. inserting the function arguments into the tool name). (2/3)

However, it achieves a six percentage point increase over Magistral Medium 1.2 (09/2025), with comparable performance on the Vals Index to GPT OSS 120B.

Full results will be posted soon.