Dec 15, 2025
Devstral 2 Models Evaluated on Coding Benchmarks
We evaluated Mistralβs new Devstral 2 models (Devstral 2 and Devstral Small 2) on our coding benchmarks (IOI, LCB, SWE-bench Verified, and Terminal-Bench). Both models are open-source and designed for software engineering tasks.
-
Devstral 2 shows strong performance on coding benchmarks, ranking 2nd among open-weight models on Terminal-Bench (43.75%) and 5th among open-weight models on SWE-bench Verified (50.4%).
-
Devstral Small 2 provides a more cost-effective option, ranking 7th among open-weight models on both Terminal-Bench (40.0%) and SWE-bench Verified (42.4%).
-
However, both models struggle on IOI and LCB, pointing to room for improvement on coding tasks.