Gemini 3.1 Pro Preview (02/26) vs Kimi K2.6: Benchmark Comparison

Gemini 3.1 Pro Preview (02/26) has the higher score on 9 of 20 shared benchmarks; Kimi K2.6 leads on 10.

The largest observed score gap is 18.92 pts on MedCode , where Gemini 3.1 Pro Preview (02/26) leads.

Reported ±1 standard-error ranges overlap on 9 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.1 Pro Preview (02/26) Kimi K2.6 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 0.00% ±0.00 1.67% ±0.83 1.67 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 20.67% ±2.81 15.87% ±2.54 4.81 pts Reported ±1 SE ranges overlap
LegalBench 87.40% ±0.33 84.74% ±0.45 2.66 pts Reported ±1 SE ranges do not overlap
EMB 52.62% ±2.97 57.85% ±2.74 5.24 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 42.98% ±1.21 44.90% ±0.73 1.92 pts Reported ±1 SE ranges overlap
MortgageTax 69.40% ±0.91 65.82% ±0.93 3.58 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 37.84% ±3.10 45.13% ±3.23 7.29 pts Reported ±1 SE ranges do not overlap
TaxEval v2 72.88% ±0.86 74.65% ±0.85 1.77 pts Reported ±1 SE ranges do not overlap
MedCode 59.06% ±2.00 40.14% ±2.04 18.92 pts Reported ±1 SE ranges do not overlap
MedScribe 76.11% ±1.92 78.15% ±1.79 2.03 pts Reported ±1 SE ranges overlap
GPQA Diamond 95.45% ±1.05 89.14% ±2.00 6.31 pts Reported ±1 SE ranges do not overlap
MMLU Pro 90.99% ±0.28 87.57% ±0.33 3.41 pts Reported ±1 SE ranges do not overlap
MMMU Pro 88.21% ±0.78 86.30% ±0.83 1.91 pts Reported ±1 SE ranges do not overlap
SAGE 48.68% ±3.29 50.22% ±3.43 1.55 pts Reported ±1 SE ranges overlap
Code Migration 17.31% ±3.93 27.77% ±4.14 10.46 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 88.48% ±0.93 86.77% ±0.97 1.71 pts Reported ±1 SE ranges overlap
ProgramBench 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
SWE-bench 78.80% ±1.83 76.20% ±1.91 2.60 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 32.03% ±4.34 37.89% ±4.91 5.86 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 53.79% ±1.30 56.63% ±1.29 2.84 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Gemini 3.1 Pro Preview (02/26) average Kimi K2.6 average
Legal 36.02% 34.09%
Finance 55.14% 57.67%
Healthcare 67.59% 59.15%
Academic 91.55% 87.67%
Education 48.68% 50.22%
Coding 43.33% 45.73%
Social Mobility 53.79% 56.63%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.1 Pro Preview (02/26) cost Kimi K2.6 cost Gemini 3.1 Pro Preview (02/26) latency Kimi K2.6 latency
Harvey's Legal Agent Benchmark $1.15 $0.82 7m17s 23m23s
Legal Research Bench $1.10 $1.52 4m02s 30m53s
LegalBench N/A N/A 10.06s 50.07s
EMB $3.77 $2.20 11m49s 39m48s
Finance Agent (v2) $1.59 $1.14 3m28s 27m10s
MortgageTax N/A N/A 23.62s 109.76s
Tax Agent Bench $0.30 $0.57 100.96s 15m24s
TaxEval v2 N/A N/A 41.93s 4m27s
MedCode N/A N/A 38.52s 5m05s
MedScribe N/A N/A 69.14s 8m14s
GPQA Diamond N/A N/A 65.76s 9m05s
MMLU Pro N/A N/A 23.91s 2m36s
MMMU Pro N/A N/A 76.99s 4m02s
SAGE N/A N/A 61.17s 4m31s
Code Migration $1.61 $5.11 12m01s 1h30m
LiveCodeBench N/A N/A 88.89s 8m49s
ProgramBench N/A N/A 12m36s 1h20m
SWE-bench $0.78 $0.27 5m12s 8m40s
Vibe Code Bench v1.1 $3.83 $1.93 20m12s 49m28s
Public Benefits Bench v1.1 $0.38 $0.91 3m06s 9m55s

Results available only for Gemini 3.1 Pro Preview (02/26)

  • Vals Index
  • ProofBench v1.1
  • BioMysteryBench
  • MysteryMechanism
  • Terminal-Bench Science
  • IOI
  • Terminal-Bench 4.0
  • Vibe Code Bench 1-100

Results available only for Kimi K2.6

None.

Model details Gemini 3.1 Pro Preview (02/26) Model details Kimi K2.6