Grok 4.6 vs GPT-6 Luna: Benchmark Comparison

Grok 4.6 has the higher score on 14 of 20 shared benchmarks; GPT-6 Luna leads on 6.

The largest observed score gap is 19.19 pts on SAGE , where GPT-6 Luna leads.

Reported ±1 standard-error ranges overlap on 9 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.6 GPT-6 Luna Gap Reported uncertainty
Vals Index 52.09% ±1.15 51.22% ±1.07 0.88 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 15.83% ±2.65 2.92% ±1.36 12.92 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 48.08% ±3.47 30.29% ±3.19 17.79 pts Reported ±1 SE ranges do not overlap
EMB 62.73% ±3.08 68.52% ±2.81 5.79 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.68% ±0.67 49.87% ±0.23 3.81 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 70.79% ±3.00 58.86% ±3.21 11.92 pts Reported ±1 SE ranges do not overlap
MedCode 44.71% ±2.26 44.69% ±2.30 0.03 pts Reported ±1 SE ranges overlap
MedScribe 86.53% ±1.96 83.71% ±1.95 2.82 pts Reported ±1 SE ranges overlap
ProofBench v1.1 51.00% ±5.02 64.00% ±4.82 13.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 72.22% ±0.00 61.48% ±2.59 10.74 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 30.63% ±3.10 19.37% ±2.66 11.26 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 11.43% ±3.83 4.29% ±2.44 7.14 pts Reported ±1 SE ranges do not overlap
SAGE 28.90% ±3.08 48.09% ±3.38 19.19 pts Reported ±1 SE ranges do not overlap
Code Migration 44.57% ±4.48 42.55% ±4.42 2.02 pts Reported ±1 SE ranges overlap
IOI 47.61% ±2.09 55.56% ±8.87 7.95 pts Reported ±1 SE ranges overlap
ProgramBench 1.00% ±0.70 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 17.17% ±1.34 13.64% ±1.51 3.54 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 76.24% ±3.82 81.65% ±3.38 5.41 pts Reported ±1 SE ranges overlap
CyberBench v1.1 67.68% ±5.87 76.25% ±5.29 8.57 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 66.85% ±1.23 57.65% ±1.28 9.20 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Grok 4.6 average GPT-6 Luna average
Legal 31.95% 16.60%
Finance 62.40% 59.08%
Healthcare 65.62% 64.20%
Math 51.00% 64.00%
Science 38.09% 28.38%
Education 28.90% 48.09%
Coding 37.32% 38.78%
Cyber 67.68% 76.25%
Social Mobility 66.85% 57.65%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.6 cost GPT-6 Luna cost Grok 4.6 latency GPT-6 Luna latency
Vals Index $4.49 $0.43 37m01s 30m34s
Harvey's Legal Agent Benchmark $4.01 $0.30 45m05s 17m32s
Legal Research Bench $1.53 $0.44 25m04s 42m28s
EMB $3.06 $0.14 52m59s 17m21s
Finance Agent (v2) $1.66 $0.12 16m41s 14m58s
Tax Agent Bench $0.98 $0.19 10m58s 33m16s
MedCode N/A N/A 2m07s 108.77s
MedScribe N/A N/A 70.94s 2m50s
ProofBench v1.1 $0.76 $0.04 14m22s 7m31s
BioMysteryBench $1.23 $0.05 14m18s 8m44s
MysteryMechanism $1.31 $0.03 27m55s 6m01s
Terminal-Bench Science $4.82 $0.24 1h18m 1h16m
SAGE N/A N/A 3m22s 91.07s
Code Migration $14.58 $0.60 1h30m 50m23s
IOI $8.18 $0.20 3h57m 35m47s
ProgramBench $16.76 $0.18 4h27m 28m55s
Terminal-Bench 4.0 $5.07 $0.35 29m27s 32m46s
Vibe Code Bench v1.1 $4.88 $1.35 25m30s 37m00s
CyberBench v1.1 $4.07 $0.13 36m52s 15m50s
Public Benefits Bench v1.1 $0.92 $0.22 25m53s 52m59s

Results available only for Grok 4.6

  • Vals RSI Index
  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
  • Time Horizon Index: KSP

Results available only for GPT-6 Luna

None.

Model details Grok 4.6 Model details GPT-6 Luna