Kimi K3 vs MiMo V2.6 Pro: Benchmark Comparison

Kimi K3 has the higher score on 9 of 18 shared benchmarks; MiMo V2.6 Pro leads on 9.

The largest observed score gap is 26.91 pts on Code Migration , where MiMo V2.6 Pro leads.

Reported ±1 standard-error ranges overlap on 10 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Kimi K3 MiMo V2.6 Pro Gap Reported uncertainty
Vals Index 50.30% ±0.99 55.20% ±1.18 4.90 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 12.92% ±2.68 10.83% ±2.45 2.08 pts Reported ±1 SE ranges overlap
Legal Research Bench 46.15% ±3.46 47.12% ±3.47 0.96 pts Reported ±1 SE ranges overlap
EMB 66.68% ±2.87 62.86% ±3.14 3.82 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.11% ±0.37 57.34% ±0.57 4.23 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 68.67% ±3.08 64.94% ±3.21 3.73 pts Reported ±1 SE ranges overlap
MedCode 49.36% ±2.20 44.97% ±2.10 4.39 pts Reported ±1 SE ranges do not overlap
MedScribe 88.05% ±1.98 88.31% ±1.94 0.26 pts Reported ±1 SE ranges overlap
ProofBench v1.1 87.00% ±3.38 70.00% ±4.61 17.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 1.43% ±1.43 2.86% ±2.01 1.43 pts Reported ±1 SE ranges overlap
SAGE 52.78% ±3.41 45.05% ±3.40 7.73 pts Reported ±1 SE ranges do not overlap
Code Migration 16.10% ±4.10 43.01% ±4.32 26.91 pts Reported ±1 SE ranges do not overlap
IOI 48.94% ±9.82 39.33% ±2.41 9.61 pts Reported ±1 SE ranges overlap
ProgramBench 2.00% ±0.99 0.50% ±0.50 1.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 17.17% ±0.51 31.31% ±3.07 14.14 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 84.97% ±2.75 85.22% ±3.39 0.25 pts Reported ±1 SE ranges overlap
CyberBench v1.1 75.24% ±5.56 72.86% ±5.50 2.38 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 68.20% ±1.21 68.94% ±1.20 0.74 pts Reported ±1 SE ranges overlap

Performance by category

Category Kimi K3 average MiMo V2.6 Pro average
Legal 29.54% 28.97%
Finance 62.82% 61.71%
Healthcare 68.70% 66.64%
Math 87.00% 70.00%
Science 1.43% 2.86%
Education 52.78% 45.05%
Coding 33.84% 39.88%
Cyber 75.24% 72.86%
Social Mobility 68.20% 68.94%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Kimi K3 cost MiMo V2.6 Pro cost Kimi K3 latency MiMo V2.6 Pro latency
Vals Index $6.38 $0.41 1h08m 1h07m
Harvey's Legal Agent Benchmark $3.83 $0.22 13m58s 25m49s
Legal Research Bench $3.47 $0.18 12m40s 30m16s
EMB $3.07 $0.33 12m21s 55m14s
Finance Agent (v2) $1.91 $0.20 4m46s 10m48s
Tax Agent Bench $2.79 $0.13 40m15s 25m48s
MedCode N/A N/A 36.99s 2m57s
MedScribe N/A N/A 38.24s 2m53s
ProofBench v1.1 $1.65 $0.25 34m31s 1h02m
Terminal-Bench Science $21.30 $0.62 4h49m 3h12m
SAGE N/A N/A 47.30s 3m40s
Code Migration $13.87 $0.69 4h18m 2h49m
IOI $14.17 $0.84 4h18m 2h26m
ProgramBench $70.48 $0.73 5h42m 3h00m
Terminal-Bench 4.0 $12.02 $0.50 3h09m 2h42m
Vibe Code Bench v1.1 $10.01 $1.04 16m39s 1h01m
CyberBench v1.1 $2.13 $0.09 30m44s 27m41s
Public Benefits Bench v1.1 $0.91 $0.08 7m19s 37m14s

Results available only for Kimi K3

  • Vals RSI Index
  • LegalBench
  • MortgageTax
  • TaxEval v2
  • BioMysteryBench
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SWE-bench
  • Vibe Code Bench 1-100
  • Time Horizon Index: KSP

Results available only for MiMo V2.6 Pro

  • MysteryMechanism
  • SRE Bench
Model details Kimi K3 Model details MiMo V2.6 Pro