Grok 4.7 vs GPT-5.6 Luna: Benchmark Comparison

Grok 4.7 has the higher score on 15 of 21 shared benchmarks; GPT-5.6 Luna leads on 6.

The largest observed score gap is 34.00 pts on ProofBench v1.1 , where GPT-5.6 Luna leads.

Reported ±1 standard-error ranges overlap on 8 of 21 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.7 GPT-5.6 Luna Gap Reported uncertainty
Vals Index 54.95% ±1.07 51.69% ±1.06 3.26 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 12.50% ±2.66 1.25% ±0.83 11.25 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 47.12% ±3.47 36.54% ±3.35 10.58 pts Reported ±1 SE ranges do not overlap
LegalBench 84.39% ±0.46 84.03% ±0.42 0.35 pts Reported ±1 SE ranges overlap
EMB 66.99% ±3.04 67.12% ±2.94 0.13 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 52.25% ±0.35 55.04% ±0.31 2.79 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 65.60% ±3.24 60.81% ±3.26 4.78 pts Reported ±1 SE ranges overlap
MedCode 49.55% ±2.17 42.39% ±2.27 7.16 pts Reported ±1 SE ranges do not overlap
MedScribe 89.38% ±1.89 84.39% ±2.58 4.99 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 26.00% ±4.41 60.00% ±4.92 34.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 69.26% ±1.33 61.48% ±0.74 7.78 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 25.23% ±2.92 14.41% ±2.36 10.81 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 10.00% ±3.61 0.00% ±0.00 10.00 pts Reported ±1 SE ranges do not overlap
SAGE 40.79% ±3.35 44.22% ±3.33 3.43 pts Reported ±1 SE ranges overlap
Code Migration 44.82% ±4.21 44.55% ±4.24 0.27 pts Reported ±1 SE ranges overlap
IOI 57.72% ±1.99 61.78% ±11.61 4.06 pts Reported ±1 SE ranges overlap
ProgramBench 0.50% ±0.50 0.00% ±0.00 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 28.79% ±2.31 11.62% ±1.01 17.17 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 86.17% ±2.18 77.06% ±3.07 9.12 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 69.46% ±5.67 73.63% ±5.56 4.17 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 65.63% ±1.24 61.16% ±1.27 4.47 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Grok 4.7 average GPT-5.6 Luna average
Legal 48.00% 40.61%
Finance 61.61% 60.99%
Healthcare 69.47% 63.39%
Math 26.00% 60.00%
Science 34.83% 25.30%
Education 40.79% 44.22%
Coding 43.60% 39.00%
Cyber 69.46% 73.63%
Social Mobility 65.63% 61.16%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.7 cost GPT-5.6 Luna cost Grok 4.7 latency GPT-5.6 Luna latency
Vals Index $12.12 $0.82 35m31s 29m27s
Harvey's Legal Agent Benchmark $11.13 $0.38 42m05s 10m56s
Legal Research Bench $4.92 $0.85 21m14s 39m34s
LegalBench N/A N/A 26.02s 5.89s
EMB $6.48 $0.37 34m57s 11m00s
Finance Agent (v2) $2.72 $0.28 15m33s 12m52s
Tax Agent Bench $1.80 $0.47 16m36s 42m46s
MedCode N/A N/A 3m33s 81.28s
MedScribe N/A N/A 2m33s 116.89s
ProofBench v1.1 $0.79 $0.12 12m47s 8m34s
BioMysteryBench $2.49 $0.09 11m37s 18m31s
MysteryMechanism $1.84 $0.11 17m33s 8m13s
Terminal-Bench Science $14.45 $0.58 1h29m 2h03m
SAGE N/A N/A 4m51s 51.88s
Code Migration $36.55 $1.88 1h09m 58m35s
IOI $12.71 $0.58 46m48s 1h10m
ProgramBench $46.49 $0.94 3h38m 41m34s
Terminal-Bench 4.0 $18.09 $0.73 50m21s 35m55s
Vibe Code Bench v1.1 $15.83 $0.73 35m47s 25m52s
CyberBench v1.1 $8.54 $0.40 15m30s 16m15s
Public Benefits Bench v1.1 $2.25 $0.27 16m57s 38m40s

Results available only for Grok 4.7

  • Vals RSI Index

Results available only for GPT-5.6 Luna

  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
Model details Grok 4.7 Model details GPT-5.6 Luna