Kimi K3 vs Hy4 Preview: Benchmark Comparison

Kimi K3 has the higher score on 14 of 19 shared benchmarks; Hy4 Preview leads on 4.

The largest observed score gap is 31.33 pts on Code Migration , where Hy4 Preview leads.

Reported ±1 standard-error ranges overlap on 9 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Kimi K3 Hy4 Preview Gap Reported uncertainty
Vals Index 50.30% ±0.99 49.94% ±1.15 0.35 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 12.92% ±2.68 9.17% ±2.00 3.75 pts Reported ±1 SE ranges overlap
Legal Research Bench 46.15% ±3.46 45.19% ±3.46 0.96 pts Reported ±1 SE ranges overlap
LegalBench 86.21% ±0.41 83.76% ±0.41 2.45 pts Reported ±1 SE ranges do not overlap
EMB 66.68% ±2.87 58.08% ±3.14 8.60 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 53.11% ±0.37 55.06% ±0.31 1.95 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 68.67% ±3.08 63.71% ±3.26 4.96 pts Reported ±1 SE ranges overlap
MedCode 49.36% ±2.20 43.25% ±2.13 6.11 pts Reported ±1 SE ranges do not overlap
MedScribe 88.05% ±1.98 83.60% ±2.06 4.45 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 87.00% ±3.38 75.00% ±4.35 12.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 72.59% ±1.61 69.26% ±2.59 3.33 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 1.43% ±1.43 1.43% ±1.43 0.00 pts Reported ±1 SE ranges overlap
Code Migration 16.10% ±4.10 47.43% ±4.27 31.33 pts Reported ±1 SE ranges do not overlap
IOI 48.94% ±9.82 59.33% ±4.60 10.39 pts Reported ±1 SE ranges overlap
ProgramBench 2.00% ±0.99 0.00% ±0.00 2.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 17.17% ±0.51 8.08% ±1.34 9.09 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 84.97% ±2.75 77.48% ±4.04 7.49 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 75.24% ±5.56 65.36% ±5.55 9.88 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 68.20% ±1.21 68.61% ±1.21 0.41 pts Reported ±1 SE ranges overlap

Performance by category

Category Kimi K3 average Hy4 Preview average
Legal 48.43% 46.04%
Finance 62.82% 58.95%
Healthcare 68.70% 63.42%
Math 87.00% 75.00%
Science 37.01% 35.34%
Coding 33.84% 38.47%
Cyber 75.24% 65.36%
Social Mobility 68.20% 68.61%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Kimi K3 cost Hy4 Preview cost Kimi K3 latency Hy4 Preview latency
Vals Index $6.38 $1.44 1h08m 1h01m
Harvey's Legal Agent Benchmark $3.83 $0.98 13m58s 29m25s
Legal Research Bench $3.47 $0.85 12m40s 1h04m
LegalBench N/A N/A 7.11s 90.56s
EMB $3.07 $0.79 12m21s 30m39s
Finance Agent (v2) $1.91 $0.59 4m46s 23m45s
Tax Agent Bench $2.79 $0.59 40m15s 41m04s
MedCode N/A N/A 36.99s 6m39s
MedScribe N/A N/A 38.24s 5m50s
ProofBench v1.1 $1.65 $0.34 34m31s 22m54s
BioMysteryBench $1.06 $0.36 4m55s 23m30s
Terminal-Bench Science $21.30 $2.14 4h49m 3h48m
Code Migration $13.87 $3.41 4h18m 2h13m
IOI $14.17 $1.35 4h18m 1h01m
ProgramBench $70.48 $12.81 5h42m 2h10m
Terminal-Bench 4.0 $12.02 $2.10 3h09m 2h10m
Vibe Code Bench v1.1 $10.01 $2.18 16m39s 41m33s
CyberBench v1.1 $2.13 $0.49 30m44s 36m43s
Public Benefits Bench v1.1 $0.91 $0.25 7m19s 1h03m

Results available only for Kimi K3

  • Vals RSI Index
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SAGE
  • LiveCodeBench
  • SWE-bench
  • Vibe Code Bench 1-100
  • Time Horizon Index: KSP

Results available only for Hy4 Preview

  • SRE Bench
Model details Kimi K3 Model details Hy4 Preview