Grok 4.5 vs GLM 5.3 Flash: Benchmark Comparison

Grok 4.5 has the higher score on 11 of 20 shared benchmarks; GLM 5.3 Flash leads on 8.

The largest observed score gap is 38.24 pts on Vibe Code Bench v1.1 , where Grok 4.5 leads.

Reported ±1 standard-error ranges overlap on 5 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.5 GLM 5.3 Flash Gap Reported uncertainty
Harvey's Legal Agent Benchmark 12.92% ±2.53 6.67% ±1.65 6.25 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 37.98% ±3.37 45.19% ±3.46 7.21 pts Reported ±1 SE ranges do not overlap
LegalBench 85.97% ±0.41 83.93% ±0.43 2.04 pts Reported ±1 SE ranges do not overlap
EMB 52.91% ±3.11 55.93% ±3.56 3.02 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 48.35% ±0.46 57.85% ±2.06 9.50 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 64.32% ±3.16 61.96% ±3.26 2.37 pts Reported ±1 SE ranges overlap
TaxEval v2 71.67% ±0.89 75.59% ±0.84 3.92 pts Reported ±1 SE ranges do not overlap
MedScribe 86.88% ±1.94 88.94% ±1.91 2.05 pts Reported ±1 SE ranges overlap
ProofBench v1.1 31.00% ±4.65 21.00% ±4.09 10.00 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 92.93% ±1.29 86.36% ±2.13 6.56 pts Reported ±1 SE ranges do not overlap
MMLU Pro 89.22% ±0.31 86.06% ±0.34 3.16 pts Reported ±1 SE ranges do not overlap
MMMU Pro 61.79% ±1.17 86.01% ±0.83 24.22 pts Reported ±1 SE ranges do not overlap
Code Migration 36.60% ±4.14 20.52% ±4.21 16.08 pts Reported ±1 SE ranges do not overlap
IOI 40.56% ±3.95 52.50% ±2.62 11.94 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 87.35% ±0.96 80.51% ±1.05 6.84 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
SkillsBench 66.03% ±4.60 40.16% ±4.43 25.88 pts Reported ±1 SE ranges do not overlap
SWE-bench 86.60% ±1.52 92.00% ±1.21 5.40 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 69.00% ±4.48 30.76% ±5.28 38.24 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 71.01% ±5.79 67.86% ±5.57 3.16 pts Reported ±1 SE ranges overlap

Performance by category

Category Grok 4.5 average GLM 5.3 Flash average
Legal 45.62% 45.26%
Finance 59.31% 62.83%
Healthcare 86.88% 88.94%
Math 31.00% 21.00%
Academic 81.31% 86.14%
Coding 55.16% 45.21%
Cyber 71.01% 67.86%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.5 cost GLM 5.3 Flash cost Grok 4.5 latency GLM 5.3 Flash latency
Harvey's Legal Agent Benchmark $2.02 $0.57 10m15s 52m17s
Legal Research Bench $1.09 $0.12 15m54s 51m07s
LegalBench N/A N/A 67.88s 12.52s
EMB $1.55 $0.71 18m28s 1h15m
Finance Agent (v2) $1.16 $0.05 4m49s 14m51s
Tax Agent Bench $0.47 $0.14 4m44s 21m37s
TaxEval v2 N/A N/A 3m15s 53.06s
MedScribe N/A N/A 11m45s 86.98s
ProofBench v1.1 $0.57 $0.22 11m09s 1h11m
GPQA Diamond N/A N/A 3m38s 2m11s
MMLU Pro N/A N/A 5m58s 44.76s
MMMU Pro N/A N/A 11m55s 28.22s
Code Migration $3.37 $3.13 23m29s 3h26m
IOI $5.64 $0.43 1h27m 1h32m
LiveCodeBench N/A N/A 9m49s 5m02s
ProgramBench $4.12 $0.80 18m59s 5h30m
SkillsBench $0.58 $0.04 2m16s 12m09s
SWE-bench $0.54 $0.02 3m20s 1h04m
Vibe Code Bench v1.1 $4.10 $2.35 13m50s 1h24m
CyberBench v1.1 $2.95 $0.14 14m08s 35m48s

Results available only for Grok 4.5

  • Vals Index
  • MortgageTax
  • MedCode
  • SAGE
  • Terminal-Bench 4.0
  • SRE Bench
  • Time Horizon Index: KSP

Results available only for GLM 5.3 Flash

  • Terminal-Bench Science
  • Vibe Code Bench 1-100
Model details Grok 4.5 Model details GLM 5.3 Flash