Grok 4.5 vs GLM 5.2: Benchmark Comparison

Grok 4.5 has the higher score on 12 of 17 shared benchmarks; GLM 5.2 leads on 5.

The largest observed score gap is 20.96 pts on SkillsBench , where Grok 4.5 leads.

Reported ±1 standard-error ranges overlap on 6 of 17 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.5 GLM 5.2 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 12.92% ±2.53 7.08% ±2.00 5.83 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 37.98% ±3.37 31.25% ±3.22 6.73 pts Reported ±1 SE ranges do not overlap
LegalBench 85.97% ±0.41 84.07% ±0.45 1.89 pts Reported ±1 SE ranges do not overlap
EMB 52.91% ±3.11 61.53% ±2.88 8.62 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 48.35% ±0.46 49.70% ±0.88 1.35 pts Reported ±1 SE ranges do not overlap
TaxEval v2 71.67% ±0.89 73.34% ±0.86 1.68 pts Reported ±1 SE ranges overlap
MedCode 43.29% ±2.31 40.77% ±2.17 2.52 pts Reported ±1 SE ranges overlap
MedScribe 86.88% ±1.94 83.53% ±2.00 3.35 pts Reported ±1 SE ranges overlap
GPQA Diamond 92.93% ±1.29 85.61% ±1.77 7.32 pts Reported ±1 SE ranges do not overlap
MMLU Pro 89.22% ±0.31 86.71% ±0.33 2.50 pts Reported ±1 SE ranges do not overlap
Code Migration 36.60% ±4.14 37.87% ±4.14 1.27 pts Reported ±1 SE ranges overlap
LiveCodeBench 87.35% ±0.96 69.50% ±1.18 17.85 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
SkillsBench 66.03% ±4.60 45.08% ±4.39 20.96 pts Reported ±1 SE ranges do not overlap
SWE-bench 86.60% ±1.52 82.80% ±1.69 3.80 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 69.00% ±4.48 63.96% ±4.79 5.04 pts Reported ±1 SE ranges overlap
SRE Bench 0.76% ±0.54 0.00% ±0.00 0.76 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Grok 4.5 average GLM 5.2 average
Legal 45.62% 40.80%
Finance 57.64% 61.52%
Healthcare 65.09% 62.15%
Academic 91.07% 86.16%
Coding 57.60% 49.95%
Cyber 0.76% 0.00%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.5 cost GLM 5.2 cost Grok 4.5 latency GLM 5.2 latency
Harvey's Legal Agent Benchmark $2.02 $2.06 10m15s 20m51s
Legal Research Bench $1.09 $0.89 15m54s 17m03s
LegalBench N/A N/A 67.88s 5.77s
EMB $1.55 $4.47 18m28s 58m52s
Finance Agent (v2) $1.16 $0.71 4m49s 7m53s
TaxEval v2 N/A N/A 3m15s 64.94s
MedCode N/A N/A 59.64s 95.27s
MedScribe N/A N/A 11m45s 2m18s
GPQA Diamond N/A N/A 3m38s 3m45s
MMLU Pro N/A N/A 5m58s 51.64s
Code Migration $3.37 $12.64 23m29s 1h54m
LiveCodeBench N/A N/A 9m49s 8m52s
ProgramBench $4.12 $12.58 18m59s 2h06m
SkillsBench $0.58 $0.81 2m16s 19m52s
SWE-bench $0.54 $0.71 3m20s 11m10s
Vibe Code Bench v1.1 $4.10 $8.49 13m50s 1h02m
SRE Bench $13.42 $26.58 1h06m 2h29m

Results available only for Grok 4.5

  • Vals Index
  • MortgageTax
  • Tax Agent Bench
  • ProofBench v1.1
  • MMMU Pro
  • SAGE
  • IOI
  • Terminal-Bench 4.0
  • CyberBench v1.1
  • Time Horizon Index: KSP

Results available only for GLM 5.2

None.

Model details Grok 4.5 Model details GLM 5.2