Kimi K3 vs GLM 5.3: Benchmark Comparison
Kimi K3 has the higher score on 12 of 25 shared benchmarks; GLM 5.3 leads on 13.
The largest observed score gap is 38.00 pts on ProofBench v1.1 , where Kimi K3 leads.
Reported ±1 standard-error ranges overlap on 9 of 24 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Kimi K3 | GLM 5.3 | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 50.30% ±0.99 | 53.51% ±1.30 | 3.22 pts | Reported ±1 SE ranges do not overlap |
| Vals RSI Index | 20.88% | 22.91% | 2.03 pts | Uncertainty comparison unavailable |
| Harvey's Legal Agent Benchmark | 12.92% ±2.68 | 8.33% ±2.00 | 4.58 pts | Reported ±1 SE ranges overlap |
| Legal Research Bench | 46.15% ±3.46 | 49.04% ±3.48 | 2.88 pts | Reported ±1 SE ranges overlap |
| LegalBench | 86.21% ±0.41 | 84.84% ±0.40 | 1.38 pts | Reported ±1 SE ranges do not overlap |
| EMB | 66.68% ±2.87 | 56.34% ±3.32 | 10.33 pts | Reported ±1 SE ranges do not overlap |
| Finance Agent (v2) | 53.11% ±0.37 | 55.84% ±2.07 | 2.73 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 68.67% ±3.08 | 73.09% ±2.97 | 4.42 pts | Reported ±1 SE ranges overlap |
| TaxEval v2 | 75.72% ±0.84 | 72.36% ±0.88 | 3.35 pts | Reported ±1 SE ranges do not overlap |
| MedCode | 49.36% ±2.20 | 42.86% ±2.11 | 6.49 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 88.05% ±1.98 | 88.81% ±2.00 | 0.76 pts | Reported ±1 SE ranges overlap |
| ProofBench v1.1 | 87.00% ±3.38 | 49.00% ±5.02 | 38.00 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench Science | 1.43% ±1.43 | 5.71% ±2.79 | 4.29 pts | Reported ±1 SE ranges do not overlap |
| GPQA Diamond | 92.93% ±1.31 | 88.13% ±1.85 | 4.80 pts | Reported ±1 SE ranges do not overlap |
| MMLU Pro | 87.97% ±0.32 | 86.77% ±0.34 | 1.20 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 16.10% ±4.10 | 44.22% ±4.29 | 28.12 pts | Reported ±1 SE ranges do not overlap |
| IOI | 48.94% ±9.82 | 68.44% ±7.48 | 19.50 pts | Reported ±1 SE ranges do not overlap |
| LiveCodeBench | 87.19% ±0.97 | 80.53% ±1.06 | 6.65 pts | Reported ±1 SE ranges do not overlap |
| ProgramBench | 2.00% ±0.99 | 1.50% ±0.86 | 0.50 pts | Reported ±1 SE ranges overlap |
| SWE-bench | 93.40% ±1.11 | 95.40% ±0.94 | 2.00 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench 4.0 | 17.17% ±0.51 | 38.89% ±1.82 | 21.72 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench 1-100 | 18.24% ±3.92 | 19.99% ±3.81 | 1.75 pts | Reported ±1 SE ranges overlap |
| Vibe Code Bench v1.1 | 84.97% ±2.75 | 78.13% ±3.94 | 6.84 pts | Reported ±1 SE ranges do not overlap |
| CyberBench v1.1 | 75.24% ±5.56 | 72.08% ±5.41 | 3.15 pts | Reported ±1 SE ranges overlap |
| Public Benefits Bench v1.1 | 68.20% ±1.21 | 68.54% ±1.21 | 0.34 pts | Reported ±1 SE ranges overlap |
Performance by category
| Category | Kimi K3 average | GLM 5.3 average |
|---|---|---|
| Legal | 48.43% | 47.40% |
| Finance | 66.04% | 64.41% |
| Healthcare | 68.70% | 65.84% |
| Math | 87.00% | 49.00% |
| Science | 1.43% | 5.71% |
| Academic | 90.45% | 87.45% |
| Coding | 46.00% | 53.39% |
| Cyber | 75.24% | 72.08% |
| Social Mobility | 68.20% | 68.54% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Kimi K3 cost | GLM 5.3 cost | Kimi K3 latency | GLM 5.3 latency |
|---|---|---|---|---|
| Vals Index | $6.38 | $7.25 | 1h08m | 1h13m |
| Vals RSI Index | $136.82 | $61.84 | 90h00m | 90h00m |
| Harvey's Legal Agent Benchmark | $3.83 | $4.30 | 13m58s | 42m38s |
| Legal Research Bench | $3.47 | $2.24 | 12m40s | 51m50s |
| LegalBench | N/A | N/A | 7.11s | 22.48s |
| EMB | $3.07 | $3.79 | 12m21s | 36m12s |
| Finance Agent (v2) | $1.91 | $1.07 | 4m46s | 15m50s |
| Tax Agent Bench | $2.79 | $1.80 | 40m15s | 50m53s |
| TaxEval v2 | N/A | N/A | 2m25s | 98.32s |
| MedCode | N/A | N/A | 36.99s | 2m49s |
| MedScribe | N/A | N/A | 38.24s | 2m02s |
| ProofBench v1.1 | $1.65 | $2.08 | 34m31s | 42m23s |
| Terminal-Bench Science | $21.30 | $15.63 | 4h49m | 2h45m |
| GPQA Diamond | N/A | N/A | 2m02s | 3m31s |
| MMLU Pro | N/A | N/A | 38.72s | 61.78s |
| Code Migration | $13.87 | $24.91 | 4h18m | 3h47m |
| IOI | $14.17 | $7.67 | 4h18m | 1h46m |
| LiveCodeBench | N/A | N/A | 3m20s | 4m07s |
| ProgramBench | $70.48 | $21.96 | 5h42m | 4h16m |
| SWE-bench | $0.76 | $0.34 | 10m19s | 14m01s |
| Terminal-Bench 4.0 | $12.02 | $9.37 | 3h09m | 1h23m |
| Vibe Code Bench 1-100 | $9.40 | $8.24 | 1h48m | 1h26m |
| Vibe Code Bench v1.1 | $10.01 | $12.45 | 16m39s | 1h04m |
| CyberBench v1.1 | $2.13 | $2.69 | 30m44s | 30m18s |
| Public Benefits Bench v1.1 | $0.91 | $0.95 | 7m19s | 35m51s |
Results available only for Kimi K3
- MortgageTax
- BioMysteryBench
- MMMU Pro
- SAGE
- Time Horizon Index: KSP
Results available only for GLM 5.3
- MysteryMechanism
- SkillsBench