GPT-6 Luna vs GLM 5.3: Benchmark Comparison

GPT-6 Luna has the higher score on 5 of 18 shared benchmarks; GLM 5.3 leads on 13.

The largest observed score gap is 25.25 pts on Terminal-Bench 4.0 , where GLM 5.3 leads.

Reported ±1 standard-error ranges overlap on 9 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark GPT-6 Luna GLM 5.3 Gap Reported uncertainty
Vals Index 51.22% ±1.07 53.51% ±1.30 2.30 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 2.92% ±1.36 8.33% ±2.00 5.42 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 30.29% ±3.19 49.04% ±3.48 18.75 pts Reported ±1 SE ranges do not overlap
EMB 68.52% ±2.81 56.34% ±3.32 12.17 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 49.87% ±0.23 55.84% ±2.07 5.97 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 58.86% ±3.21 73.09% ±2.97 14.23 pts Reported ±1 SE ranges do not overlap
MedCode 44.69% ±2.30 42.86% ±2.11 1.82 pts Reported ±1 SE ranges overlap
MedScribe 83.71% ±1.95 88.81% ±2.00 5.10 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 64.00% ±4.82 49.00% ±5.02 15.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 19.37% ±2.66 22.97% ±2.83 3.60 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 4.29% ±2.44 5.71% ±2.79 1.43 pts Reported ±1 SE ranges overlap
Code Migration 42.55% ±4.42 44.22% ±4.29 1.67 pts Reported ±1 SE ranges overlap
IOI 55.56% ±8.87 68.44% ±7.48 12.89 pts Reported ±1 SE ranges overlap
ProgramBench 0.50% ±0.50 1.50% ±0.86 1.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 13.64% ±1.51 38.89% ±1.82 25.25 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 81.65% ±3.38 78.13% ±3.94 3.52 pts Reported ±1 SE ranges overlap
CyberBench v1.1 76.25% ±5.29 72.08% ±5.41 4.17 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 57.65% ±1.28 68.54% ±1.21 10.89 pts Reported ±1 SE ranges do not overlap

Performance by category

Category GPT-6 Luna average GLM 5.3 average
Legal 16.60% 28.69%
Finance 59.08% 61.76%
Healthcare 64.20% 65.84%
Math 64.00% 49.00%
Science 11.83% 14.34%
Coding 38.78% 46.24%
Cyber 76.25% 72.08%
Social Mobility 57.65% 68.54%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark GPT-6 Luna cost GLM 5.3 cost GPT-6 Luna latency GLM 5.3 latency
Vals Index $0.43 $7.25 30m34s 1h13m
Harvey's Legal Agent Benchmark $0.30 $4.30 17m32s 42m38s
Legal Research Bench $0.44 $2.24 42m28s 51m50s
EMB $0.14 $3.79 17m21s 36m12s
Finance Agent (v2) $0.12 $1.07 14m58s 15m50s
Tax Agent Bench $0.19 $1.80 33m16s 50m53s
MedCode N/A N/A 108.77s 2m49s
MedScribe N/A N/A 2m50s 2m02s
ProofBench v1.1 $0.04 $2.08 7m31s 42m23s
MysteryMechanism $0.03 $1.53 6m01s 30m42s
Terminal-Bench Science $0.24 $15.63 1h16m 2h45m
Code Migration $0.60 $24.91 50m23s 3h47m
IOI $0.20 $7.67 35m47s 1h46m
ProgramBench $0.18 $21.96 28m55s 4h16m
Terminal-Bench 4.0 $0.35 $9.37 32m46s 1h23m
Vibe Code Bench v1.1 $1.35 $12.45 37m00s 1h04m
CyberBench v1.1 $0.13 $2.69 15m50s 30m18s
Public Benefits Bench v1.1 $0.22 $0.95 52m59s 35m51s

Results available only for GPT-6 Luna

  • BioMysteryBench
  • SAGE

Results available only for GLM 5.3

  • Vals RSI Index
  • LegalBench
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
Model details GPT-6 Luna Model details GLM 5.3