Release Date: Jul 11, 2025

Developer Moonshot AI πŸ‡¨πŸ‡³
Context Window 128k
Max Output Tokens 16k
Token Costs (in/out) $1.00/3.00
Weights Open
Input Modalities

Accuracy

57.15 %

Avg. Cost (In/Out)

$ 1.00 / $ 3.00

Latency

5 min 42 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.90
89/139

0.0%

Β±2.27
87/132

0.0%

Β±0.63
59/62

0.0%

Β±1.16
81/136

0.0%

Β±0.44
69/136

0.0%

Β±0.39
93/132
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Moonshot AI
Temperature: 0.3
Top P: Default
Top K: Default
Max Output Tokens: 16,384

Updates

Jul 30, 2025

Our SWE-bench Verified evaluation of Kimi K2 Instruct achieved 34% accuracy, barely more than half of Kimi’s published figures!

After investigating the model responses, we identified the two following sources of error:

  1. The model struggles to use tools - it often includes tool calls in the response itself! We replicated the issue on multiple popular inference providers. However, even discarding such errors only increases accuracy by around 2%.
  2. The model often gets stuck repeating itself, leading to unnecessarily long and incorrect responses. This is a common failure mode of models at zero temperature, though it’s most prevalent among thinking models.

Jul 22, 2025

We found that Kimi K2 Instruct is the new state-of-the-art open-source model according to our evaluations.

The model cracks the top 10 on Math500 and LiveCodeBench, narrowly beating out DeepSeek R1 on both. On other public benchmarks, however, Kimi K2 Instruct delivers middle-of-the-pack performance.

However, Kimi K2 Instruct struggles on our proprietary benchmarks, failing to break the top 10 on any of them. We noticed it particularly struggles with legal tasks such as Case Law and Contract Law but performs comparatively better on finance tasks such as Corp Fin and Tax Eval.

The model offers solid value at 1.00input/1.00 input/3.00 output per million tokens, which is cheaper than DeepSeek R1 (3.00/3.00/7.00, both as hosted on Together AI) but more expensive than Mistral Medium 3.1 (05/2025) (0.40/0.40/2.00).

We’re currently evaluating the model on SWE-bench Verified, on which Kimi’s reported accuracy would top our leaderboard. Looking forward to seeing whether the model can live up to the hype!