Kimi K3 vs GLM 5.3: Benchmark Comparison

Kimi K3 has the higher score on 12 of 25 shared benchmarks; GLM 5.3 leads on 13.

The largest observed score gap is 38.00 pts on ProofBench v1.1 , where Kimi K3 leads.

Reported ±1 standard-error ranges overlap on 9 of 24 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Kimi K3 GLM 5.3 Gap Reported uncertainty
Vals Index 50.30% ±0.99 53.51% ±1.30 3.22 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 20.88% 22.91% 2.03 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 12.92% ±2.68 8.33% ±2.00 4.58 pts Reported ±1 SE ranges overlap
Legal Research Bench 46.15% ±3.46 49.04% ±3.48 2.88 pts Reported ±1 SE ranges overlap
LegalBench 86.21% ±0.41 84.84% ±0.40 1.38 pts Reported ±1 SE ranges do not overlap
EMB 66.68% ±2.87 56.34% ±3.32 10.33 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 53.11% ±0.37 55.84% ±2.07 2.73 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 68.67% ±3.08 73.09% ±2.97 4.42 pts Reported ±1 SE ranges overlap
TaxEval v2 75.72% ±0.84 72.36% ±0.88 3.35 pts Reported ±1 SE ranges do not overlap
MedCode 49.36% ±2.20 42.86% ±2.11 6.49 pts Reported ±1 SE ranges do not overlap
MedScribe 88.05% ±1.98 88.81% ±2.00 0.76 pts Reported ±1 SE ranges overlap
ProofBench v1.1 87.00% ±3.38 49.00% ±5.02 38.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 1.43% ±1.43 5.71% ±2.79 4.29 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 92.93% ±1.31 88.13% ±1.85 4.80 pts Reported ±1 SE ranges do not overlap
MMLU Pro 87.97% ±0.32 86.77% ±0.34 1.20 pts Reported ±1 SE ranges do not overlap
Code Migration 16.10% ±4.10 44.22% ±4.29 28.12 pts Reported ±1 SE ranges do not overlap
IOI 48.94% ±9.82 68.44% ±7.48 19.50 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 87.19% ±0.97 80.53% ±1.06 6.65 pts Reported ±1 SE ranges do not overlap
ProgramBench 2.00% ±0.99 1.50% ±0.86 0.50 pts Reported ±1 SE ranges overlap
SWE-bench 93.40% ±1.11 95.40% ±0.94 2.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 17.17% ±0.51 38.89% ±1.82 21.72 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 18.24% ±3.92 19.99% ±3.81 1.75 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 84.97% ±2.75 78.13% ±3.94 6.84 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 75.24% ±5.56 72.08% ±5.41 3.15 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 68.20% ±1.21 68.54% ±1.21 0.34 pts Reported ±1 SE ranges overlap

Performance by category

Category Kimi K3 average GLM 5.3 average
Legal 48.43% 47.40%
Finance 66.04% 64.41%
Healthcare 68.70% 65.84%
Math 87.00% 49.00%
Science 1.43% 5.71%
Academic 90.45% 87.45%
Coding 46.00% 53.39%
Cyber 75.24% 72.08%
Social Mobility 68.20% 68.54%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Kimi K3 cost GLM 5.3 cost Kimi K3 latency GLM 5.3 latency
Vals Index $6.38 $7.25 1h08m 1h13m
Vals RSI Index $136.82 $61.84 90h00m 90h00m
Harvey's Legal Agent Benchmark $3.83 $4.30 13m58s 42m38s
Legal Research Bench $3.47 $2.24 12m40s 51m50s
LegalBench N/A N/A 7.11s 22.48s
EMB $3.07 $3.79 12m21s 36m12s
Finance Agent (v2) $1.91 $1.07 4m46s 15m50s
Tax Agent Bench $2.79 $1.80 40m15s 50m53s
TaxEval v2 N/A N/A 2m25s 98.32s
MedCode N/A N/A 36.99s 2m49s
MedScribe N/A N/A 38.24s 2m02s
ProofBench v1.1 $1.65 $2.08 34m31s 42m23s
Terminal-Bench Science $21.30 $15.63 4h49m 2h45m
GPQA Diamond N/A N/A 2m02s 3m31s
MMLU Pro N/A N/A 38.72s 61.78s
Code Migration $13.87 $24.91 4h18m 3h47m
IOI $14.17 $7.67 4h18m 1h46m
LiveCodeBench N/A N/A 3m20s 4m07s
ProgramBench $70.48 $21.96 5h42m 4h16m
SWE-bench $0.76 $0.34 10m19s 14m01s
Terminal-Bench 4.0 $12.02 $9.37 3h09m 1h23m
Vibe Code Bench 1-100 $9.40 $8.24 1h48m 1h26m
Vibe Code Bench v1.1 $10.01 $12.45 16m39s 1h04m
CyberBench v1.1 $2.13 $2.69 30m44s 30m18s
Public Benefits Bench v1.1 $0.91 $0.95 7m19s 35m51s

Results available only for Kimi K3

  • MortgageTax
  • BioMysteryBench
  • MMMU Pro
  • SAGE
  • Time Horizon Index: KSP

Results available only for GLM 5.3

  • MysteryMechanism
  • SkillsBench
Model details Kimi K3 Model details GLM 5.3