Jan 27, 2026
Kimi K2.5 Evaluated on (almost) all benchmarks!
Weβve finished our evaluation of Kimi K2.5 and continue to find excellent performance across the board. Here are some highlights:
- The model places first on CorpFin by a significant margin (2%), though it takes substantially longer to respond than other models in the top 10.
- The model is remarkably consistent - its worst placement is 16th, on CaseLaw, on which it remains within 10% of first place.
- The model is head and shoulders above its open-weight competition, placing first among open-weight models on 13 of the 17 benchmarks on which it was evaluated.
We ran all benchmarks on the native provider with temperature 1 as recommended by Kimi, with 128K max tokens on all but the coding benchmarks, for which we used 256K.
Stay tuned for our evaluation on Vibe Code Bench and congrats again to Moonshot on the release!