Kimi K2.5

Kimi K2.5 is an open-weight model from Moonshot AI, released January 26, 2026. Its best result is #19 of 90 on SAGE.

Release Date: Jan 26, 2026

Developer Moonshot AI Β πŸ‡¨πŸ‡³
Context Window 262k
Max Output Tokens 128k
Token Costs (in/out) $0.60/3.00
Weights Open
Input Modalities

Accuracy

47.81 %

Avg. Cost (In/Out)

$ 0.60 / $ 3.00

Latency

299 min 7 s

Vals Index
BenchmarksAccuracyRankings

6.85%

24/24

6.96%

Β±1.29
62/73

28.47%

Β±2.23
60/70

35.79%

Β±1.99
63/74

15.87%

Β±2.54
54/73

39.32%

Β±2.12
66/104

76.44%

Β±1.99
73/106

66.53%

Β±0.93
34/98

49.87%

Β±3.36
19/90

74.20%

Β±0.85
40/145

17.54%

Β±3.26
81/108

84.09%

Β±2.13
55/138

83.87%

Β±1.03
46/143

85.91%

Β±0.34
52/138

84.33%

Β±0.87
26/93

70.00%

Β±2.05
61/88

41.95%

Β±0.75
67/75
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Moonshot AI
Temperature: 1
Top P: 0.95
Top K: Default
Max Output Tokens: 128,000

Updates

Jan 27, 2026

We’ve finished our evaluation of Kimi K2.5 and continue to find excellent performance across the board. Here are some highlights:

  • The model places first on CorpFin by a significant margin (2%), though it takes substantially longer to respond than other models in the top 10.
  • The model is remarkably consistent - its worst placement is 16th, on CaseLaw, on which it remains within 10% of first place.
  • The model is head and shoulders above its open-weight competition, placing first among open-weight models on 13 of the 17 benchmarks on which it was evaluated.

We ran all benchmarks on the native provider with temperature 1 as recommended by Kimi, with 128K max tokens on all but the coding benchmarks, for which we used 256K.

Stay tuned for our evaluation on Vibe Code Bench and congrats again to Moonshot on the release!

Jan 26, 2026

Kimi K2.5 is the new #1 open-weight model, taking the top spot on both our Vals Index and our Vals Multimodal Index. Here are the key takeaways:

  • Like its predecessor Kimi K2 Thinking, the model excels across coding tasks, placing first across open-weight models on both SWE-bench Verified and Terminal-Bench 2.0.
  • Unlike its predecessor, it also excels on our Finance Agent and SAGE, placing third across all models on the latter. By comparison, Kimi K2 Thinking didn’t even support multimodal input!
  • Aside from being open-weight, another significant differentiation between Kimi and other top models is price - it’s the cheapest model in the top 10 of both Vals indices!

We ran all benchmarks on the native provider with temperature 1 as recommended by Kimi, with 128K max tokens on all but the coding benchmarks, for which we used 256K.