Kimi K3 vs Hy4 Preview: Benchmark Comparison
Kimi K3 has the higher score on 14 of 19 shared benchmarks; Hy4 Preview leads on 4.
The largest observed score gap is 31.33 pts on Code Migration , where Hy4 Preview leads.
Reported ±1 standard-error ranges overlap on 9 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Kimi K3 | Hy4 Preview | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 50.30% ±0.99 | 49.94% ±1.15 | 0.35 pts | Reported ±1 SE ranges overlap |
| Harvey's Legal Agent Benchmark | 12.92% ±2.68 | 9.17% ±2.00 | 3.75 pts | Reported ±1 SE ranges overlap |
| Legal Research Bench | 46.15% ±3.46 | 45.19% ±3.46 | 0.96 pts | Reported ±1 SE ranges overlap |
| LegalBench | 86.21% ±0.41 | 83.76% ±0.41 | 2.45 pts | Reported ±1 SE ranges do not overlap |
| EMB | 66.68% ±2.87 | 58.08% ±3.14 | 8.60 pts | Reported ±1 SE ranges do not overlap |
| Finance Agent (v2) | 53.11% ±0.37 | 55.06% ±0.31 | 1.95 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 68.67% ±3.08 | 63.71% ±3.26 | 4.96 pts | Reported ±1 SE ranges overlap |
| MedCode | 49.36% ±2.20 | 43.25% ±2.13 | 6.11 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 88.05% ±1.98 | 83.60% ±2.06 | 4.45 pts | Reported ±1 SE ranges do not overlap |
| ProofBench v1.1 | 87.00% ±3.38 | 75.00% ±4.35 | 12.00 pts | Reported ±1 SE ranges do not overlap |
| BioMysteryBench | 72.59% ±1.61 | 69.26% ±2.59 | 3.33 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench Science | 1.43% ±1.43 | 1.43% ±1.43 | 0.00 pts | Reported ±1 SE ranges overlap |
| Code Migration | 16.10% ±4.10 | 47.43% ±4.27 | 31.33 pts | Reported ±1 SE ranges do not overlap |
| IOI | 48.94% ±9.82 | 59.33% ±4.60 | 10.39 pts | Reported ±1 SE ranges overlap |
| ProgramBench | 2.00% ±0.99 | 0.00% ±0.00 | 2.00 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench 4.0 | 17.17% ±0.51 | 8.08% ±1.34 | 9.09 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 84.97% ±2.75 | 77.48% ±4.04 | 7.49 pts | Reported ±1 SE ranges do not overlap |
| CyberBench v1.1 | 75.24% ±5.56 | 65.36% ±5.55 | 9.88 pts | Reported ±1 SE ranges overlap |
| Public Benefits Bench v1.1 | 68.20% ±1.21 | 68.61% ±1.21 | 0.41 pts | Reported ±1 SE ranges overlap |
Performance by category
| Category | Kimi K3 average | Hy4 Preview average |
|---|---|---|
| Legal | 48.43% | 46.04% |
| Finance | 62.82% | 58.95% |
| Healthcare | 68.70% | 63.42% |
| Math | 87.00% | 75.00% |
| Science | 37.01% | 35.34% |
| Coding | 33.84% | 38.47% |
| Cyber | 75.24% | 65.36% |
| Social Mobility | 68.20% | 68.61% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Kimi K3 cost | Hy4 Preview cost | Kimi K3 latency | Hy4 Preview latency |
|---|---|---|---|---|
| Vals Index | $6.38 | $1.44 | 1h08m | 1h01m |
| Harvey's Legal Agent Benchmark | $3.83 | $0.98 | 13m58s | 29m25s |
| Legal Research Bench | $3.47 | $0.85 | 12m40s | 1h04m |
| LegalBench | N/A | N/A | 7.11s | 90.56s |
| EMB | $3.07 | $0.79 | 12m21s | 30m39s |
| Finance Agent (v2) | $1.91 | $0.59 | 4m46s | 23m45s |
| Tax Agent Bench | $2.79 | $0.59 | 40m15s | 41m04s |
| MedCode | N/A | N/A | 36.99s | 6m39s |
| MedScribe | N/A | N/A | 38.24s | 5m50s |
| ProofBench v1.1 | $1.65 | $0.34 | 34m31s | 22m54s |
| BioMysteryBench | $1.06 | $0.36 | 4m55s | 23m30s |
| Terminal-Bench Science | $21.30 | $2.14 | 4h49m | 3h48m |
| Code Migration | $13.87 | $3.41 | 4h18m | 2h13m |
| IOI | $14.17 | $1.35 | 4h18m | 1h01m |
| ProgramBench | $70.48 | $12.81 | 5h42m | 2h10m |
| Terminal-Bench 4.0 | $12.02 | $2.10 | 3h09m | 2h10m |
| Vibe Code Bench v1.1 | $10.01 | $2.18 | 16m39s | 41m33s |
| CyberBench v1.1 | $2.13 | $0.49 | 30m44s | 36m43s |
| Public Benefits Bench v1.1 | $0.91 | $0.25 | 7m19s | 1h03m |
Results available only for Kimi K3
- Vals RSI Index
- MortgageTax
- TaxEval v2
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- SAGE
- LiveCodeBench
- SWE-bench
- Vibe Code Bench 1-100
- Time Horizon Index: KSP
Results available only for Hy4 Preview
- SRE Bench