Grok 4.7 vs GPT-5.6 Terra: Benchmark Comparison

Grok 4.7 has the higher score on 10 of 19 shared benchmarks; GPT-5.6 Terra leads on 7.

The largest observed score gap is 48.00 pts on ProofBench v1.1 , where GPT-5.6 Terra leads.

Reported ±1 standard-error ranges overlap on 11 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.7 GPT-5.6 Terra Gap Reported uncertainty
Vals Index 54.95% ±1.07 53.09% ±1.29 1.86 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 12.50% ±2.66 0.83% ±0.83 11.67 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 47.12% ±3.47 41.35% ±3.42 5.77 pts Reported ±1 SE ranges overlap
LegalBench 84.39% ±0.46 85.11% ±0.45 0.72 pts Reported ±1 SE ranges overlap
EMB 66.99% ±3.04 66.20% ±3.11 0.79 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 52.25% ±0.35 54.44% ±2.07 2.19 pts Reported ±1 SE ranges overlap
Tax Agent Bench 65.60% ±3.24 65.20% ±3.19 0.40 pts Reported ±1 SE ranges overlap
MedCode 49.55% ±2.17 43.41% ±2.17 6.14 pts Reported ±1 SE ranges do not overlap
MedScribe 89.38% ±1.89 82.87% ±1.95 6.51 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 26.00% ±4.41 74.00% ±4.41 48.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 10.00% ±3.61 10.00% ±3.61 0.00 pts Reported ±1 SE ranges overlap
SAGE 40.79% ±3.35 47.00% ±3.40 6.21 pts Reported ±1 SE ranges overlap
Code Migration 44.82% ±4.21 47.80% ±4.28 2.99 pts Reported ±1 SE ranges overlap
IOI 57.72% ±1.99 87.61% ±6.53 29.89 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 0.50% ±0.50 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 28.79% ±2.31 22.73% ±2.31 6.06 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 86.17% ±2.18 74.59% ±4.20 11.58 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 69.46% ±5.67 72.08% ±5.41 2.62 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 65.63% ±1.24 62.38% ±1.26 3.25 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Grok 4.7 average GPT-5.6 Terra average
Legal 48.00% 42.43%
Finance 61.61% 61.95%
Healthcare 69.47% 63.14%
Math 26.00% 74.00%
Science 10.00% 10.00%
Education 40.79% 47.00%
Coding 43.60% 46.65%
Cyber 69.46% 72.08%
Social Mobility 65.63% 62.38%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.7 cost GPT-5.6 Terra cost Grok 4.7 latency GPT-5.6 Terra latency
Vals Index $12.12 $5.78 35m31s 35m39s
Harvey's Legal Agent Benchmark $11.13 $3.20 42m05s 17m05s
Legal Research Bench $4.92 $7.91 21m14s 1h10m
LegalBench N/A N/A 26.02s 2.78s
EMB $6.48 $2.19 34m57s 13m33s
Finance Agent (v2) $2.72 $3.63 15m33s 24m10s
Tax Agent Bench $1.80 $4.83 16m36s 1h01m
MedCode N/A N/A 3m33s 18.41s
MedScribe N/A N/A 2m33s 35.53s
ProofBench v1.1 $0.79 $1.27 12m47s 12m57s
Terminal-Bench Science $14.45 $5.19 1h29m 2h04m
SAGE N/A N/A 4m51s 40.06s
Code Migration $36.55 $8.13 1h09m 48m34s
IOI $12.71 $8.66 46m48s 1h23m
ProgramBench $46.49 $5.38 3h38m 33m36s
Terminal-Bench 4.0 $18.09 $5.60 50m21s 31m29s
Vibe Code Bench v1.1 $15.83 $7.89 35m47s 19m12s
CyberBench v1.1 $8.54 $3.31 15m30s 17m04s
Public Benefits Bench v1.1 $2.25 $1.20 16m57s 19m31s

Results available only for Grok 4.7

  • Vals RSI Index
  • BioMysteryBench
  • MysteryMechanism

Results available only for GPT-5.6 Terra

  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
Model details Grok 4.7 Model details GPT-5.6 Terra