Grok 4.6 vs Kimi K3: Benchmark Comparison

Grok 4.6 has the higher score on 13 of 29 shared benchmarks; Kimi K3 leads on 15.

The largest observed score gap is 36.00 pts on ProofBench v1.1 , where Kimi K3 leads.

Reported ±1 standard-error ranges overlap on 18 of 28 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.6 Kimi K3 Gap Reported uncertainty
Vals Index 52.09% ±1.15 50.30% ±0.99 1.80 pts Reported ±1 SE ranges overlap
Vals RSI Index 25.46% 20.88% 4.58 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 15.83% ±2.65 12.92% ±2.68 2.92 pts Reported ±1 SE ranges overlap
Legal Research Bench 48.08% ±3.47 46.15% ±3.46 1.92 pts Reported ±1 SE ranges overlap
LegalBench 86.31% ±0.42 86.21% ±0.41 0.09 pts Reported ±1 SE ranges overlap
EMB 62.73% ±3.08 66.68% ±2.87 3.95 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.68% ±0.67 53.11% ±0.37 0.57 pts Reported ±1 SE ranges overlap
MortgageTax 64.19% ±0.95 66.34% ±0.91 2.15 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 70.79% ±3.00 68.67% ±3.08 2.12 pts Reported ±1 SE ranges overlap
TaxEval v2 71.10% ±0.90 75.72% ±0.84 4.62 pts Reported ±1 SE ranges do not overlap
MedCode 44.71% ±2.26 49.36% ±2.20 4.64 pts Reported ±1 SE ranges do not overlap
MedScribe 86.53% ±1.96 88.05% ±1.98 1.51 pts Reported ±1 SE ranges overlap
ProofBench v1.1 51.00% ±5.02 87.00% ±3.38 36.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 72.22% ±0.00 72.59% ±1.61 0.37 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 11.43% ±3.83 1.43% ±1.43 10.00 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 94.70% ±1.13 92.93% ±1.31 1.77 pts Reported ±1 SE ranges overlap
MMLU Pro 89.40% ±0.30 87.97% ±0.32 1.43 pts Reported ±1 SE ranges do not overlap
SAGE 28.90% ±3.08 52.78% ±3.41 23.88 pts Reported ±1 SE ranges do not overlap
Code Migration 44.57% ±4.48 16.10% ±4.10 28.47 pts Reported ±1 SE ranges do not overlap
IOI 47.61% ±2.09 48.94% ±9.82 1.33 pts Reported ±1 SE ranges overlap
LiveCodeBench 88.22% ±0.94 87.19% ±0.97 1.04 pts Reported ±1 SE ranges overlap
ProgramBench 1.00% ±0.70 2.00% ±0.99 1.00 pts Reported ±1 SE ranges overlap
SWE-bench 95.60% ±0.92 93.40% ±1.11 2.20 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 17.17% ±1.34 17.17% ±0.51 0.00 pts Reported ±1 SE ranges overlap
Vibe Code Bench 1-100 14.75% ±3.17 18.24% ±3.92 3.49 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 76.24% ±3.82 84.97% ±2.75 8.73 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 67.68% ±5.87 75.24% ±5.56 7.56 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 66.85% ±1.23 68.20% ±1.21 1.35 pts Reported ±1 SE ranges overlap
Time Horizon Index: KSP 5.83% ±0.00 10.50% ±5.38 4.67 pts Reported ±1 SE ranges overlap

Performance by category

Category Grok 4.6 average Kimi K3 average
Legal 50.07% 48.43%
Finance 64.50% 66.10%
Healthcare 65.62% 68.70%
Math 51.00% 87.00%
Science 41.83% 37.01%
Academic 92.05% 90.45%
Education 28.90% 52.78%
Coding 48.15% 46.00%
Cyber 67.68% 75.24%
Social Mobility 66.85% 68.20%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.6 cost Kimi K3 cost Grok 4.6 latency Kimi K3 latency
Vals Index $4.49 $6.38 37m01s 1h08m
Vals RSI Index $88.12 $136.82 90h00m 90h00m
Harvey's Legal Agent Benchmark $4.01 $3.83 45m05s 13m58s
Legal Research Bench $1.53 $3.47 25m04s 12m40s
LegalBench N/A N/A 27.42s 7.11s
EMB $3.06 $3.07 52m59s 12m21s
Finance Agent (v2) $1.66 $1.91 16m41s 4m46s
MortgageTax N/A N/A 8.55s 78.94s
Tax Agent Bench $0.98 $2.79 10m58s 40m15s
TaxEval v2 N/A N/A 47.57s 2m25s
MedCode N/A N/A 2m07s 36.99s
MedScribe N/A N/A 70.94s 38.24s
ProofBench v1.1 $0.76 $1.65 14m22s 34m31s
BioMysteryBench $1.23 $1.06 14m18s 4m55s
Terminal-Bench Science $4.82 $21.30 1h18m 4h49m
GPQA Diamond N/A N/A 3m29s 2m02s
MMLU Pro N/A N/A 54.98s 38.72s
SAGE N/A N/A 3m22s 47.30s
Code Migration $14.58 $13.87 1h30m 4h18m
IOI $8.18 $14.17 3h57m 4h18m
LiveCodeBench N/A N/A 118.40s 3m20s
ProgramBench $16.76 $70.48 4h27m 5h42m
SWE-bench $0.78 $0.76 10m04s 10m19s
Terminal-Bench 4.0 $5.07 $12.02 29m27s 3h09m
Vibe Code Bench 1-100 $26.27 $9.40 2h11m 1h48m
Vibe Code Bench v1.1 $4.88 $10.01 25m30s 16m39s
CyberBench v1.1 $4.07 $2.13 36m52s 30m44s
Public Benefits Bench v1.1 $0.92 $0.91 25m53s 7m19s
Time Horizon Index: KSP $225.12 $216.43 N/A N/A

Results available only for Grok 4.6

  • MysteryMechanism
  • SkillsBench

Results available only for Kimi K3

  • MMMU Pro
Model details Grok 4.6 Model details Kimi K3