Claude Opus 5 vs GPT-5.6 Terra: Benchmark Comparison

Claude Opus 5 has the higher score on 25 of 28 shared benchmarks; GPT-5.6 Terra leads on 3.

The largest observed score gap is 30.81 pts on Terminal-Bench 4.0 , where Claude Opus 5 leads.

Reported ±1 standard-error ranges overlap on 7 of 28 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Opus 5 GPT-5.6 Terra Gap Reported uncertainty
Vals Index 63.67% ±0.96 53.09% ±1.29 10.59 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 6.67% ±1.65 0.83% ±0.83 5.83 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 55.29% ±3.46 41.35% ±3.42 13.94 pts Reported ±1 SE ranges do not overlap
LegalBench 86.97% ±0.42 85.11% ±0.45 1.87 pts Reported ±1 SE ranges do not overlap
EMB 73.56% ±2.24 66.20% ±3.11 7.35 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 58.63% ±0.08 54.44% ±2.07 4.20 pts Reported ±1 SE ranges do not overlap
MortgageTax 72.06% ±0.88 67.33% ±0.93 4.73 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 75.06% ±2.96 65.20% ±3.19 9.86 pts Reported ±1 SE ranges do not overlap
TaxEval v2 75.14% ±0.83 76.17% ±0.84 1.02 pts Reported ±1 SE ranges overlap
MedCode 63.57% ±1.99 43.41% ±2.17 20.16 pts Reported ±1 SE ranges do not overlap
MedScribe 90.98% ±1.92 82.87% ±1.95 8.12 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 99.00% ±1.00 74.00% ±4.41 25.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 27.14% ±5.35 10.00% ±3.61 17.14 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 93.43% ±1.24 90.91% ±1.91 2.52 pts Reported ±1 SE ranges overlap
MMLU Pro 91.59% ±0.28 86.66% ±0.33 4.93 pts Reported ±1 SE ranges do not overlap
MMMU Pro 89.88% ±0.72 86.47% ±0.82 3.41 pts Reported ±1 SE ranges do not overlap
SAGE 49.43% ±3.29 47.00% ±3.40 2.42 pts Reported ±1 SE ranges overlap
Code Migration 57.47% ±4.37 47.80% ±4.28 9.67 pts Reported ±1 SE ranges do not overlap
IOI 84.33% ±9.96 87.61% ±6.53 3.28 pts Reported ±1 SE ranges overlap
LiveCodeBench 89.03% ±0.91 85.93% ±1.02 3.10 pts Reported ±1 SE ranges do not overlap
ProgramBench 3.00% ±1.21 0.50% ±0.50 2.50 pts Reported ±1 SE ranges do not overlap
SkillsBench 60.44% ±4.58 58.90% ±4.47 1.55 pts Reported ±1 SE ranges overlap
SWE-bench 97.00% ±0.76 95.40% ±0.94 1.60 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 53.53% ±1.34 22.73% ±2.31 30.81 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 28.53% ±4.33 14.82% ±3.12 13.71 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 88.40% ±3.00 74.59% ±4.20 13.81 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 65.36% ±5.55 72.08% ±5.41 6.73 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 76.93% ±1.10 62.38% ±1.26 14.55 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Opus 5 average GPT-5.6 Terra average
Legal 49.64% 42.43%
Finance 70.89% 65.87%
Healthcare 77.28% 63.14%
Math 99.00% 74.00%
Science 27.14% 10.00%
Academic 91.64% 88.01%
Education 49.43% 47.00%
Coding 62.42% 54.25%
Cyber 65.36% 72.08%
Social Mobility 76.93% 62.38%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Opus 5 cost GPT-5.6 Terra cost Claude Opus 5 latency GPT-5.6 Terra latency
Vals Index $19.31 $5.78 57m41s 35m39s
Harvey's Legal Agent Benchmark $23.67 $3.20 55m37s 17m05s
Legal Research Bench $6.76 $7.91 31m01s 1h10m
LegalBench N/A N/A 4.42s 2.78s
EMB $6.07 $2.19 25m59s 13m33s
Finance Agent (v2) $5.12 $3.63 9m58s 24m10s
MortgageTax N/A N/A 7.98s 6.32s
Tax Agent Bench $5.14 $4.83 16m48s 1h01m
TaxEval v2 N/A N/A 58.34s 37.28s
MedCode N/A N/A 24.50s 18.41s
MedScribe N/A N/A 76.56s 35.53s
ProofBench v1.1 $2.11 $1.27 9m51s 12m57s
Terminal-Bench Science $32.54 $5.19 2h12m 2h04m
GPQA Diamond N/A N/A 39.59s 36.76s
MMLU Pro N/A N/A 9.60s 8.17s
MMMU Pro N/A N/A 30.03s 13.03s
SAGE N/A N/A 56.17s 40.06s
Code Migration $60.51 $8.13 2h51m 48m34s
IOI $16.48 $8.66 56m42s 1h23m
LiveCodeBench N/A N/A 59.22s 44.33s
ProgramBench $60.29 $5.38 2h55m 33m36s
SkillsBench $2.51 $1.75 8m25s 6m55s
SWE-bench $1.29 $0.40 9m37s 3m00s
Terminal-Bench 4.0 $18.60 $5.60 1h07m 31m29s
Vibe Code Bench 1-100 $41.49 $33.33 1h22m 1h49m
Vibe Code Bench v1.1 $33.88 $7.89 1h28m 19m12s
CyberBench v1.1 $2.87 $3.31 19m25s 17m04s
Public Benefits Bench v1.1 $3.95 $1.20 30m09s 19m31s

Results available only for Claude Opus 5

  • Vals RSI Index
  • BioMysteryBench
  • MysteryMechanism
  • SRE Bench
  • CUA-bench
  • Time Horizon Index: KSP

Results available only for GPT-5.6 Terra

None.

Model details Claude Opus 5 Model details GPT-5.6 Terra