Kimi K3 vs GPT-6 Luna: Benchmark Comparison

Kimi K3 has the higher score on 13 of 19 shared benchmarks; GPT-6 Luna leads on 6.

The largest observed score gap is 26.46 pts on Code Migration , where GPT-6 Luna leads.

Reported ±1 standard-error ranges overlap on 7 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Kimi K3 GPT-6 Luna Gap Reported uncertainty
Vals Index 50.30% ±0.99 51.22% ±1.07 0.92 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 12.92% ±2.68 2.92% ±1.36 10.00 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 46.15% ±3.46 30.29% ±3.19 15.87 pts Reported ±1 SE ranges do not overlap
EMB 66.68% ±2.87 68.52% ±2.81 1.84 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.11% ±0.37 49.87% ±0.23 3.24 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 68.67% ±3.08 58.86% ±3.21 9.81 pts Reported ±1 SE ranges do not overlap
MedCode 49.36% ±2.20 44.69% ±2.30 4.67 pts Reported ±1 SE ranges do not overlap
MedScribe 88.05% ±1.98 83.71% ±1.95 4.34 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 87.00% ±3.38 64.00% ±4.82 23.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 72.59% ±1.61 61.48% ±2.59 11.11 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 1.43% ±1.43 4.29% ±2.44 2.86 pts Reported ±1 SE ranges overlap
SAGE 52.78% ±3.41 48.09% ±3.38 4.69 pts Reported ±1 SE ranges overlap
Code Migration 16.10% ±4.10 42.55% ±4.42 26.46 pts Reported ±1 SE ranges do not overlap
IOI 48.94% ±9.82 55.56% ±8.87 6.61 pts Reported ±1 SE ranges overlap
ProgramBench 2.00% ±0.99 0.50% ±0.50 1.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 17.17% ±0.51 13.64% ±1.51 3.54 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 84.97% ±2.75 81.65% ±3.38 3.32 pts Reported ±1 SE ranges overlap
CyberBench v1.1 75.24% ±5.56 76.25% ±5.29 1.01 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 68.20% ±1.21 57.65% ±1.28 10.55 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Kimi K3 average GPT-6 Luna average
Legal 29.54% 16.60%
Finance 62.82% 59.08%
Healthcare 68.70% 64.20%
Math 87.00% 64.00%
Science 37.01% 32.88%
Education 52.78% 48.09%
Coding 33.84% 38.78%
Cyber 75.24% 76.25%
Social Mobility 68.20% 57.65%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Kimi K3 cost GPT-6 Luna cost Kimi K3 latency GPT-6 Luna latency
Vals Index $6.38 $0.43 1h08m 30m34s
Harvey's Legal Agent Benchmark $3.83 $0.30 13m58s 17m32s
Legal Research Bench $3.47 $0.44 12m40s 42m28s
EMB $3.07 $0.14 12m21s 17m21s
Finance Agent (v2) $1.91 $0.12 4m46s 14m58s
Tax Agent Bench $2.79 $0.19 40m15s 33m16s
MedCode N/A N/A 36.99s 108.77s
MedScribe N/A N/A 38.24s 2m50s
ProofBench v1.1 $1.65 $0.04 34m31s 7m31s
BioMysteryBench $1.06 $0.05 4m55s 8m44s
Terminal-Bench Science $21.30 $0.24 4h49m 1h16m
SAGE N/A N/A 47.30s 91.07s
Code Migration $13.87 $0.60 4h18m 50m23s
IOI $14.17 $0.20 4h18m 35m47s
ProgramBench $70.48 $0.18 5h42m 28m55s
Terminal-Bench 4.0 $12.02 $0.35 3h09m 32m46s
Vibe Code Bench v1.1 $10.01 $1.35 16m39s 37m00s
CyberBench v1.1 $2.13 $0.13 30m44s 15m50s
Public Benefits Bench v1.1 $0.91 $0.22 7m19s 52m59s

Results available only for Kimi K3

  • Vals RSI Index
  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SWE-bench
  • Vibe Code Bench 1-100
  • Time Horizon Index: KSP

Results available only for GPT-6 Luna

  • MysteryMechanism
Model details Kimi K3 Model details GPT-6 Luna