Grok 4.6 vs GPT-5.6 Luna: Benchmark Comparison

Grok 4.6 has the higher score on 17 of 28 shared benchmarks; GPT-5.6 Luna leads on 11.

The largest observed score gap is 16.22 pts on MysteryMechanism , where Grok 4.6 leads.

Reported ±1 standard-error ranges overlap on 9 of 28 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.6 GPT-5.6 Luna Gap Reported uncertainty
Vals Index 52.09% ±1.15 51.69% ±1.06 0.41 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 15.83% ±2.65 1.25% ±0.83 14.58 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 48.08% ±3.47 36.54% ±3.35 11.54 pts Reported ±1 SE ranges do not overlap
LegalBench 86.31% ±0.42 84.03% ±0.42 2.27 pts Reported ±1 SE ranges do not overlap
EMB 62.73% ±3.08 67.12% ±2.94 4.39 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.68% ±0.67 55.04% ±0.31 1.36 pts Reported ±1 SE ranges do not overlap
MortgageTax 64.19% ±0.95 67.29% ±0.92 3.10 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 70.79% ±3.00 60.81% ±3.26 9.97 pts Reported ±1 SE ranges do not overlap
TaxEval v2 71.10% ±0.90 76.17% ±0.84 5.07 pts Reported ±1 SE ranges do not overlap
MedCode 44.71% ±2.26 42.39% ±2.27 2.32 pts Reported ±1 SE ranges overlap
MedScribe 86.53% ±1.96 84.39% ±2.58 2.14 pts Reported ±1 SE ranges overlap
ProofBench v1.1 51.00% ±5.02 60.00% ±4.92 9.00 pts Reported ±1 SE ranges overlap
BioMysteryBench 72.22% ±0.00 61.48% ±0.74 10.74 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 30.63% ±3.10 14.41% ±2.36 16.22 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 11.43% ±3.83 0.00% ±0.00 11.43 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 94.70% ±1.13 91.67% ±1.74 3.03 pts Reported ±1 SE ranges do not overlap
MMLU Pro 89.40% ±0.30 86.04% ±0.35 3.36 pts Reported ±1 SE ranges do not overlap
SAGE 28.90% ±3.08 44.22% ±3.33 15.32 pts Reported ±1 SE ranges do not overlap
Code Migration 44.57% ±4.48 44.55% ±4.24 0.02 pts Reported ±1 SE ranges overlap
IOI 47.61% ±2.09 61.78% ±11.61 14.17 pts Reported ±1 SE ranges do not overlap
ProgramBench 1.00% ±0.70 0.00% ±0.00 1.00 pts Reported ±1 SE ranges do not overlap
SkillsBench 55.77% ±4.84 60.45% ±4.75 4.68 pts Reported ±1 SE ranges overlap
SWE-bench 95.60% ±0.92 93.00% ±1.14 2.60 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 17.17% ±1.34 11.62% ±1.01 5.56 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 14.75% ±3.17 22.59% ±4.00 7.84 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 76.24% ±3.82 77.06% ±3.07 0.82 pts Reported ±1 SE ranges overlap
CyberBench v1.1 67.68% ±5.87 73.63% ±5.56 5.95 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 66.85% ±1.23 61.16% ±1.27 5.68 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Grok 4.6 average GPT-5.6 Luna average
Legal 50.07% 40.61%
Finance 64.50% 65.29%
Healthcare 65.62% 63.39%
Math 51.00% 60.00%
Science 38.09% 25.30%
Academic 92.05% 88.85%
Education 28.90% 44.22%
Coding 44.09% 46.38%
Cyber 67.68% 73.63%
Social Mobility 66.85% 61.16%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.6 cost GPT-5.6 Luna cost Grok 4.6 latency GPT-5.6 Luna latency
Vals Index $4.49 $0.82 37m01s 29m27s
Harvey's Legal Agent Benchmark $4.01 $0.38 45m05s 10m56s
Legal Research Bench $1.53 $0.85 25m04s 39m34s
LegalBench N/A N/A 27.42s 5.89s
EMB $3.06 $0.37 52m59s 11m00s
Finance Agent (v2) $1.66 $0.28 16m41s 12m52s
MortgageTax N/A N/A 8.55s 22.84s
Tax Agent Bench $0.98 $0.47 10m58s 42m46s
TaxEval v2 N/A N/A 47.57s 76.52s
MedCode N/A N/A 2m07s 81.28s
MedScribe N/A N/A 70.94s 116.89s
ProofBench v1.1 $0.76 $0.12 14m22s 8m34s
BioMysteryBench $1.23 $0.09 14m18s 18m31s
MysteryMechanism $1.31 $0.11 27m55s 8m13s
Terminal-Bench Science $4.82 $0.58 1h18m 2h03m
GPQA Diamond N/A N/A 3m29s 53.52s
MMLU Pro N/A N/A 54.98s 16.95s
SAGE N/A N/A 3m22s 51.88s
Code Migration $14.58 $1.88 1h30m 58m35s
IOI $8.18 $0.58 3h57m 1h10m
ProgramBench $16.76 $0.94 4h27m 41m34s
SkillsBench $0.92 $0.21 13m28s 5m20s
SWE-bench $0.78 $0.04 10m04s 3m21s
Terminal-Bench 4.0 $5.07 $0.73 29m27s 35m55s
Vibe Code Bench 1-100 $26.27 $4.91 2h11m 2h03m
Vibe Code Bench v1.1 $4.88 $0.73 25m30s 25m52s
CyberBench v1.1 $4.07 $0.40 36m52s 16m15s
Public Benefits Bench v1.1 $0.92 $0.27 25m53s 38m40s

Results available only for Grok 4.6

  • Vals RSI Index
  • LiveCodeBench
  • Time Horizon Index: KSP

Results available only for GPT-5.6 Luna

  • MMMU Pro
Model details Grok 4.6 Model details GPT-5.6 Luna