Model comparison

Gemini 3.7 Flash vs Kimi K3: Benchmark Comparison

Gemini 3.7 Flash has the higher score on 13 of 26 shared benchmarks; Kimi K3 leads on 13.

The largest observed score gap is 31.49 pts on CyberBench v1.1 , where Kimi K3 leads.

Reported ±1 standard-error ranges overlap on 10 of 25 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.7 Flash Kimi K3 Gap Reported uncertainty
Vals Index 59.31% ±1.06 57.81% ±1.06 1.49 pts Reported ±1 SE ranges overlap
Vals RSI Index 18.27% 20.88% 2.61 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 8.75% ±2.15 10.83% ±2.29 2.08 pts Reported ±1 SE ranges overlap
Legal Research Bench 34.62% ±3.31 44.23% ±3.45 9.62 pts Reported ±1 SE ranges do not overlap
LegalBench 87.26% ±0.42 86.02% ±0.39 1.24 pts Reported ±1 SE ranges do not overlap
EMB 71.33% ±2.26 66.40% ±3.04 4.93 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 59.04% ±0.27 54.36% ±0.53 4.68 pts Reported ±1 SE ranges do not overlap
MortgageTax 66.65% ±0.92 66.34% ±0.91 0.32 pts Reported ±1 SE ranges overlap
Tax Agent Bench 57.66% ±3.31 68.67% ±3.08 11.01 pts Reported ±1 SE ranges do not overlap
TaxEval v2 74.73% ±0.85 75.72% ±0.84 0.98 pts Reported ±1 SE ranges overlap
MedCode 53.39% ±2.12 48.88% ±2.19 4.51 pts Reported ±1 SE ranges do not overlap
MedScribe 83.94% ±2.00 87.96% ±1.89 4.02 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 58.00% ±4.96 87.00% ±3.38 29.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 5.71% ±2.79 2.86% ±2.01 2.86 pts Reported ±1 SE ranges overlap
GPQA Diamond 93.94% ±1.49 92.93% ±1.31 1.01 pts Reported ±1 SE ranges overlap
MMLU Pro 90.12% ±0.30 87.97% ±0.32 2.15 pts Reported ±1 SE ranges do not overlap
MMMU Pro 88.96% ±0.75 88.15% ±0.78 0.81 pts Reported ±1 SE ranges overlap
SAGE 49.23% ±3.38 54.26% ±3.42 5.03 pts Reported ±1 SE ranges overlap
Code Migration 34.80% ±4.22 16.10% ±4.10 18.70 pts Reported ±1 SE ranges do not overlap
IOI 67.83% ±3.61 48.94% ±9.82 18.89 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 88.65% ±0.92 87.19% ±0.97 1.47 pts Reported ±1 SE ranges overlap
ProgramBench 0.00% ±0.00 2.00% ±0.99 2.00 pts Reported ±1 SE ranges do not overlap
SWE-bench 80.80% ±1.76 93.40% ±1.11 12.60 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 6.06% ±0.88 12.63% ±1.01 6.56 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 70.39% ±4.84 84.96% ±2.69 14.57 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 43.75% ±2.21 75.24% ±5.56 31.49 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Gemini 3.7 Flash average Kimi K3 average
Index 38.79% 39.35%
Legal 43.54% 47.03%
Finance 65.88% 66.30%
Healthcare 68.67% 68.42%
Math 58.00% 87.00%
Science 5.71% 2.86%
Academic 91.01% 89.68%
Education 49.23% 54.26%
Coding 49.79% 49.32%
Beta 43.75% 75.24%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.7 Flash cost Kimi K3 cost Gemini 3.7 Flash latency Kimi K3 latency
Vals Index $4.17 $6.47 43m13s 1h10m
Vals RSI Index $180.16 $136.82 90h00m 90h00m
Harvey's Legal Agent Benchmark $2.56 $5.62 7m32s 47m23s
Legal Research Bench $0.76 $3.22 4m49s 35m46s
LegalBench N/A N/A 2.23s 18.52s
EMB $5.69 $3.48 10m51s 34m53s
Finance Agent (v2) $1.48 $1.07 3m01s 14m37s
MortgageTax N/A N/A 4.11s 78.94s
Tax Agent Bench $0.48 $2.79 87.24s 40m15s
TaxEval v2 N/A N/A 7.83s 2m25s
MedCode N/A N/A 9.16s 116.18s
MedScribe N/A N/A 16.82s 2m17s
ProofBench v1.1 $0.56 $1.67 7m04s 34m31s
Terminal-Bench Science N/A N/A 1h08m 4h28m
GPQA Diamond N/A N/A 8.77s 2m02s
MMLU Pro N/A N/A 4.16s 38.72s
MMMU Pro N/A N/A 6.28s 106.52s
SAGE N/A N/A 9.79s 2m29s
Code Migration $21.46 $13.87 2h39m 4h18m
IOI $3.40 $14.17 21m09s 4h18m
LiveCodeBench N/A N/A 14.53s 3m20s
ProgramBench $7.13 $70.48 21m30s 5h42m
SWE-bench $1.44 $0.76 5m50s 10m19s
Terminal-Bench 4.0 $14.07 $8.07 1h59m 3h13m
Vibe Code Bench v1.1 $4.83 $17.59 23m43s 1h27m
CyberBench v1.1 $2.15 $1.48 7m34s 30m44s

Results available only for Gemini 3.7 Flash

  • SkillsBench
  • SRE Bench

Results available only for Kimi K3

  • BioMysteryBench
  • Vibe Code Bench 1-100
  • Time Horizon Index: KSP
  • Public Benefits Bench v1.1
Model details Gemini 3.7 Flash Model details Kimi K3