Grok 4.6 vs GPT-5.6 Terra: Benchmark Comparison

Grok 4.6 has the higher score on 14 of 27 shared benchmarks; GPT-5.6 Terra leads on 13.

The largest observed score gap is 40.00 pts on IOI , where GPT-5.6 Terra leads.

Reported ±1 standard-error ranges overlap on 15 of 27 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.6 GPT-5.6 Terra Gap Reported uncertainty
Vals Index 52.09% ±1.15 53.09% ±1.29 0.99 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 15.83% ±2.65 0.83% ±0.83 15.00 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 48.08% ±3.47 41.35% ±3.42 6.73 pts Reported ±1 SE ranges overlap
LegalBench 86.31% ±0.42 85.11% ±0.45 1.20 pts Reported ±1 SE ranges do not overlap
EMB 62.73% ±3.08 66.20% ±3.11 3.48 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.68% ±0.67 54.44% ±2.07 0.75 pts Reported ±1 SE ranges overlap
MortgageTax 64.19% ±0.95 67.33% ±0.93 3.14 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 70.79% ±3.00 65.20% ±3.19 5.59 pts Reported ±1 SE ranges overlap
TaxEval v2 71.10% ±0.90 76.17% ±0.84 5.07 pts Reported ±1 SE ranges do not overlap
MedCode 44.71% ±2.26 43.41% ±2.17 1.30 pts Reported ±1 SE ranges overlap
MedScribe 86.53% ±1.96 82.87% ±1.95 3.67 pts Reported ±1 SE ranges overlap
ProofBench v1.1 51.00% ±5.02 74.00% ±4.41 23.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 11.43% ±3.83 10.00% ±3.61 1.43 pts Reported ±1 SE ranges overlap
GPQA Diamond 94.70% ±1.13 90.91% ±1.91 3.79 pts Reported ±1 SE ranges do not overlap
MMLU Pro 89.40% ±0.30 86.66% ±0.33 2.74 pts Reported ±1 SE ranges do not overlap
SAGE 28.90% ±3.08 47.00% ±3.40 18.10 pts Reported ±1 SE ranges do not overlap
Code Migration 44.57% ±4.48 47.80% ±4.28 3.23 pts Reported ±1 SE ranges overlap
IOI 47.61% ±2.09 87.61% ±6.53 40.00 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 88.22% ±0.94 85.93% ±1.02 2.29 pts Reported ±1 SE ranges do not overlap
ProgramBench 1.00% ±0.70 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
SkillsBench 55.77% ±4.84 58.90% ±4.47 3.13 pts Reported ±1 SE ranges overlap
SWE-bench 95.60% ±0.92 95.40% ±0.94 0.20 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 17.17% ±1.34 22.73% ±2.31 5.55 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 14.75% ±3.17 14.82% ±3.12 0.07 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 76.24% ±3.82 74.59% ±4.20 1.65 pts Reported ±1 SE ranges overlap
CyberBench v1.1 67.68% ±5.87 72.08% ±5.41 4.41 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 66.85% ±1.23 62.38% ±1.26 4.46 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Grok 4.6 average GPT-5.6 Terra average
Legal 50.07% 42.43%
Finance 64.50% 65.87%
Healthcare 65.62% 63.14%
Math 51.00% 74.00%
Science 11.43% 10.00%
Academic 92.05% 88.78%
Education 28.90% 47.00%
Coding 48.99% 54.25%
Cyber 67.68% 72.08%
Social Mobility 66.85% 62.38%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.6 cost GPT-5.6 Terra cost Grok 4.6 latency GPT-5.6 Terra latency
Vals Index $4.49 $5.78 37m01s 35m39s
Harvey's Legal Agent Benchmark $4.01 $3.20 45m05s 17m05s
Legal Research Bench $1.53 $7.91 25m04s 1h10m
LegalBench N/A N/A 27.42s 2.78s
EMB $3.06 $2.19 52m59s 13m33s
Finance Agent (v2) $1.66 $3.63 16m41s 24m10s
MortgageTax N/A N/A 8.55s 6.32s
Tax Agent Bench $0.98 $4.83 10m58s 1h01m
TaxEval v2 N/A N/A 47.57s 37.28s
MedCode N/A N/A 2m07s 18.41s
MedScribe N/A N/A 70.94s 35.53s
ProofBench v1.1 $0.76 $1.27 14m22s 12m57s
Terminal-Bench Science $4.82 $5.19 1h18m 2h04m
GPQA Diamond N/A N/A 3m29s 36.76s
MMLU Pro N/A N/A 54.98s 8.17s
SAGE N/A N/A 3m22s 40.06s
Code Migration $14.58 $8.13 1h30m 48m34s
IOI $8.18 $8.66 3h57m 1h23m
LiveCodeBench N/A N/A 118.40s 44.33s
ProgramBench $16.76 $5.38 4h27m 33m36s
SkillsBench $0.92 $1.75 13m28s 6m55s
SWE-bench $0.78 $0.40 10m04s 3m00s
Terminal-Bench 4.0 $5.07 $5.60 29m27s 31m29s
Vibe Code Bench 1-100 $26.27 $33.33 2h11m 1h49m
Vibe Code Bench v1.1 $4.88 $7.89 25m30s 19m12s
CyberBench v1.1 $4.07 $3.31 36m52s 17m04s
Public Benefits Bench v1.1 $0.92 $1.20 25m53s 19m31s

Results available only for Grok 4.6

  • Vals RSI Index
  • BioMysteryBench
  • MysteryMechanism
  • Time Horizon Index: KSP

Results available only for GPT-5.6 Terra

  • MMMU Pro
Model details Grok 4.6 Model details GPT-5.6 Terra