Claude Sonnet 5 vs Kimi K3: Benchmark Comparison

Claude Sonnet 5 has the higher score on 5 of 27 shared benchmarks; Kimi K3 leads on 22.

The largest observed score gap is 28.29 pts on Code Migration , where Claude Sonnet 5 leads.

Reported ±1 standard-error ranges overlap on 12 of 27 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5 Kimi K3 Gap Reported uncertainty
Vals Index 51.77% ±1.09 50.30% ±0.99 1.48 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 5.00% ±1.65 12.92% ±2.68 7.92 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.83% ±3.43 46.15% ±3.46 4.33 pts Reported ±1 SE ranges overlap
LegalBench 83.92% ±0.46 86.21% ±0.41 2.29 pts Reported ±1 SE ranges do not overlap
EMB 66.32% ±3.01 66.68% ±2.87 0.36 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.91% ±0.52 53.11% ±0.37 0.80 pts Reported ±1 SE ranges overlap
MortgageTax 70.03% ±0.90 66.34% ±0.91 3.70 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.27% ±3.19 68.67% ±3.08 6.40 pts Reported ±1 SE ranges do not overlap
TaxEval v2 75.63% ±0.84 75.72% ±0.84 0.08 pts Reported ±1 SE ranges overlap
MedCode 47.54% ±2.27 49.36% ±2.20 1.82 pts Reported ±1 SE ranges overlap
MedScribe 76.05% ±3.05 88.05% ±1.98 11.99 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 77.00% ±4.23 87.00% ±3.38 10.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 5.71% ±2.79 1.43% ±1.43 4.29 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 88.89% ±2.22 92.93% ±1.31 4.04 pts Reported ±1 SE ranges do not overlap
MMLU Pro 87.55% ±0.37 87.97% ±0.32 0.43 pts Reported ±1 SE ranges overlap
MMMU Pro 83.01% ±0.90 88.15% ±0.78 5.14 pts Reported ±1 SE ranges do not overlap
SAGE 48.92% ±3.40 52.78% ±3.41 3.86 pts Reported ±1 SE ranges overlap
Code Migration 44.39% ±4.25 16.10% ±4.10 28.29 pts Reported ±1 SE ranges do not overlap
IOI 45.00% ±2.75 48.94% ±9.82 3.94 pts Reported ±1 SE ranges overlap
LiveCodeBench 82.43% ±1.09 87.19% ±0.97 4.76 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 2.00% ±0.99 2.00 pts Reported ±1 SE ranges do not overlap
SWE-bench 79.60% ±1.80 93.40% ±1.11 13.80 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 9.60% ±1.01 17.17% ±0.51 7.58 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 13.82% ±2.98 18.24% ±3.92 4.42 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 81.33% ±3.05 84.97% ±2.75 3.64 pts Reported ±1 SE ranges overlap
CyberBench v1.1 61.91% ±5.74 75.24% ±5.56 13.33 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 66.03% ±1.23 68.20% ±1.21 2.17 pts Reported ±1 SE ranges overlap

Performance by category

Category Claude Sonnet 5 average Kimi K3 average
Legal 43.58% 48.43%
Finance 65.63% 66.10%
Healthcare 61.80% 68.70%
Math 77.00% 87.00%
Science 5.71% 1.43%
Academic 86.48% 89.68%
Education 48.92% 52.78%
Coding 44.52% 46.00%
Cyber 61.91% 75.24%
Social Mobility 66.03% 68.20%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5 cost Kimi K3 cost Claude Sonnet 5 latency Kimi K3 latency
Vals Index $13.72 $6.38 54m16s 1h08m
Harvey's Legal Agent Benchmark $8.95 $3.83 38m52s 13m58s
Legal Research Bench $2.72 $3.47 25m46s 12m40s
LegalBench N/A N/A 4.90s 7.11s
EMB $10.29 $3.07 49m46s 12m21s
Finance Agent (v2) $0.75 $1.91 13m12s 4m46s
MortgageTax N/A N/A 28.67s 78.94s
Tax Agent Bench $1.79 $2.79 19m19s 40m15s
TaxEval v2 N/A N/A 3m22s 2m25s
MedCode N/A N/A 2m15s 36.99s
MedScribe N/A N/A 4m12s 38.24s
ProofBench v1.1 $1.37 $1.65 14m58s 34m31s
Terminal-Bench Science $28.28 $21.30 2h59m 4h49m
GPQA Diamond N/A N/A 63.23s 2m02s
MMLU Pro N/A N/A 25.19s 38.72s
MMMU Pro N/A N/A 18.82s 106.52s
SAGE N/A N/A 7m14s 47.30s
Code Migration $35.31 $13.87 1h57m 4h18m
IOI $12.89 $14.17 59m44s 4h18m
LiveCodeBench N/A N/A 77.02s 3m20s
ProgramBench $24.36 $70.48 1h31m 5h42m
SWE-bench $1.49 $0.76 16m02s 10m19s
Terminal-Bench 4.0 $26.33 $12.02 1h45m 3h09m
Vibe Code Bench 1-100 $71.15 $9.40 4h16m 1h48m
Vibe Code Bench v1.1 $25.39 $10.01 1h07m 16m39s
CyberBench v1.1 $1.89 $2.13 16m57s 30m44s
Public Benefits Bench v1.1 $1.29 $0.91 21m09s 7m19s

Results available only for Claude Sonnet 5

  • SkillsBench

Results available only for Kimi K3

  • Vals RSI Index
  • BioMysteryBench
  • Time Horizon Index: KSP
Model details Claude Sonnet 5 Model details Kimi K3