Kimi K2.6 vs MiniMax-M3: Benchmark Comparison

Kimi K2.6 has the higher score on 7 of 19 shared benchmarks; MiniMax-M3 leads on 12.

The largest observed score gap is 13.94 pts on Legal Research Bench , where MiniMax-M3 leads.

Reported ±1 standard-error ranges overlap on 6 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Kimi K2.6 MiniMax-M3 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 1.67% ±0.83 4.17% ±1.65 2.50 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 15.87% ±2.54 29.81% ±3.18 13.94 pts Reported ±1 SE ranges do not overlap
LegalBench 84.74% ±0.45 85.42% ±0.44 0.68 pts Reported ±1 SE ranges overlap
EMB 57.85% ±2.74 47.76% ±2.95 10.10 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 44.90% ±0.73 48.27% ±0.44 3.37 pts Reported ±1 SE ranges do not overlap
MortgageTax 65.82% ±0.93 68.36% ±0.91 2.54 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 45.13% ±3.23 49.69% ±3.25 4.56 pts Reported ±1 SE ranges overlap
TaxEval v2 74.65% ±0.85 72.73% ±0.86 1.92 pts Reported ±1 SE ranges do not overlap
MedCode 40.14% ±2.04 46.29% ±2.10 6.15 pts Reported ±1 SE ranges do not overlap
MedScribe 78.15% ±1.79 87.25% ±1.96 9.10 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 89.14% ±2.00 92.68% ±1.44 3.53 pts Reported ±1 SE ranges do not overlap
MMLU Pro 87.57% ±0.33 84.22% ±0.36 3.35 pts Reported ±1 SE ranges do not overlap
MMMU Pro 86.30% ±0.83 81.16% ±0.94 5.14 pts Reported ±1 SE ranges do not overlap
SAGE 50.22% ±3.43 50.57% ±3.44 0.35 pts Reported ±1 SE ranges overlap
Code Migration 27.77% ±4.14 19.93% ±3.94 7.84 pts Reported ±1 SE ranges overlap
LiveCodeBench 86.77% ±0.97 82.15% ±1.05 4.62 pts Reported ±1 SE ranges do not overlap
SWE-bench 76.20% ±1.91 75.00% ±1.94 1.20 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 37.89% ±4.91 47.57% ±5.44 9.68 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 56.63% ±1.29 64.14% ±1.25 7.51 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Kimi K2.6 average MiniMax-M3 average
Legal 34.09% 39.80%
Finance 57.67% 57.36%
Healthcare 59.15% 66.77%
Academic 87.67% 86.02%
Education 50.22% 50.57%
Coding 57.16% 56.16%
Social Mobility 56.63% 64.14%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Kimi K2.6 cost MiniMax-M3 cost Kimi K2.6 latency MiniMax-M3 latency
Harvey's Legal Agent Benchmark $0.82 $1.46 23m23s 22m31s
Legal Research Bench $1.52 $0.34 30m53s 13m34s
LegalBench N/A N/A 50.07s 8.21s
EMB $2.20 $2.09 39m48s 31m30s
Finance Agent (v2) $1.14 $0.32 27m10s 8m17s
MortgageTax N/A N/A 109.76s 26.76s
Tax Agent Bench $0.57 $0.16 15m24s 5m15s
TaxEval v2 N/A N/A 4m27s 94.04s
MedCode N/A N/A 5m05s 63.12s
MedScribe N/A N/A 8m14s 2m04s
GPQA Diamond N/A N/A 9m05s 4m39s
MMLU Pro N/A N/A 2m36s 41.39s
MMMU Pro N/A N/A 4m02s 68.35s
SAGE N/A N/A 4m31s 118.54s
Code Migration $5.11 $7.07 1h30m 1h14m
LiveCodeBench N/A N/A 8m49s 6m07s
SWE-bench $0.27 $0.42 8m40s 12m06s
Vibe Code Bench v1.1 $1.93 $6.45 49m28s 1h25m
Public Benefits Bench v1.1 $0.91 $0.25 9m55s 11m31s

Results available only for Kimi K2.6

  • ProgramBench

Results available only for MiniMax-M3

  • Vals Index
  • ProofBench v1.1
  • SkillsBench
  • Terminal-Bench 4.0
  • Vibe Code Bench 1-100
  • CyberBench v1.1
Model details Kimi K2.6 Model details MiniMax-M3