Model comparison

Claude Sonnet 5 vs GPT-5.6 Terra: Benchmark Comparison

Claude Sonnet 5 has the higher score on 10 of 28 shared benchmarks; GPT-5.6 Terra leads on 18.

The largest observed score gap is 42.61 pts on IOI , where GPT-5.6 Terra leads.

Reported ±1 standard-error ranges overlap on 15 of 28 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5 GPT-5.6 Terra Gap Reported uncertainty
Vals Index 51.77% ±1.09 53.09% ±1.29 1.31 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 5.00% ±1.65 0.83% ±0.83 4.17 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.83% ±3.43 41.35% ±3.42 0.48 pts Reported ±1 SE ranges overlap
LegalBench 83.92% ±0.46 85.11% ±0.45 1.18 pts Reported ±1 SE ranges do not overlap
EMB 66.32% ±3.01 66.20% ±3.11 0.11 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.91% ±0.52 54.44% ±2.07 0.53 pts Reported ±1 SE ranges overlap
MortgageTax 70.03% ±0.90 67.33% ±0.93 2.70 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.27% ±3.19 65.20% ±3.19 2.93 pts Reported ±1 SE ranges overlap
TaxEval v2 75.63% ±0.84 76.17% ±0.84 0.53 pts Reported ±1 SE ranges overlap
MedCode 47.54% ±2.27 43.41% ±2.17 4.13 pts Reported ±1 SE ranges overlap
MedScribe 76.05% ±3.05 82.87% ±1.95 6.81 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 77.00% ±4.23 74.00% ±4.41 3.00 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 2.86% ±2.01 10.00% ±3.61 7.14 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 88.89% ±2.22 90.91% ±1.91 2.02 pts Reported ±1 SE ranges overlap
MMLU Pro 87.55% ±0.37 86.66% ±0.33 0.89 pts Reported ±1 SE ranges do not overlap
MMMU Pro 83.01% ±0.90 86.47% ±0.82 3.47 pts Reported ±1 SE ranges do not overlap
SAGE 48.92% ±3.40 47.00% ±3.40 1.92 pts Reported ±1 SE ranges overlap
Code Migration 44.39% ±4.25 47.80% ±4.28 3.41 pts Reported ±1 SE ranges overlap
IOI 45.00% ±2.75 87.61% ±6.53 42.61 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 82.43% ±1.09 85.93% ±1.02 3.50 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
SkillsBench 46.48% ±4.49 58.90% ±4.47 12.42 pts Reported ±1 SE ranges do not overlap
SWE-bench 79.60% ±1.80 95.40% ±0.94 15.80 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 9.60% ±1.01 22.73% ±2.31 13.13 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 13.82% ±2.98 14.82% ±3.12 1.00 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 81.33% ±3.05 74.59% ±4.20 6.73 pts Reported ±1 SE ranges overlap
CyberBench v1.1 61.91% ±5.74 72.08% ±5.41 10.18 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 66.03% ±1.23 62.38% ±1.26 3.65 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Sonnet 5 average GPT-5.6 Terra average
Legal 43.58% 42.43%
Finance 65.63% 65.87%
Healthcare 61.80% 63.14%
Math 77.00% 74.00%
Science 2.86% 10.00%
Academic 86.48% 88.01%
Education 48.92% 47.00%
Coding 44.74% 54.25%
Beta 61.91% 72.08%
Social Mobility 66.03% 62.38%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5 cost GPT-5.6 Terra cost Claude Sonnet 5 latency GPT-5.6 Terra latency
Vals Index $13.72 $5.78 54m16s 35m39s
Harvey's Legal Agent Benchmark $8.95 $3.20 38m52s 17m05s
Legal Research Bench $2.72 $7.91 25m46s 1h10m
LegalBench N/A N/A 4.90s 2.78s
EMB $10.29 $2.19 49m46s 13m33s
Finance Agent (v2) $0.75 $3.63 13m12s 24m10s
MortgageTax N/A N/A 28.67s 6.32s
Tax Agent Bench $1.79 $4.83 19m19s 1h01m
TaxEval v2 N/A N/A 3m22s 37.28s
MedCode N/A N/A 2m15s 18.41s
MedScribe N/A N/A 4m12s 35.53s
ProofBench v1.1 $1.37 $1.27 14m58s 12m57s
Terminal-Bench Science $18.70 $5.19 2h55m 2h04m
GPQA Diamond N/A N/A 63.23s 36.76s
MMLU Pro N/A N/A 25.19s 8.17s
MMMU Pro N/A N/A 18.82s 13.03s
SAGE N/A N/A 7m14s 40.06s
Code Migration $35.31 $8.13 1h57m 48m34s
IOI $12.89 $8.66 59m44s 1h23m
LiveCodeBench N/A N/A 77.02s 44.33s
ProgramBench $24.36 $5.38 1h31m 33m36s
SkillsBench $2.90 $1.75 15m12s 6m55s
SWE-bench $1.49 $0.40 16m02s 3m00s
Terminal-Bench 4.0 $26.33 $5.60 1h45m 31m29s
Vibe Code Bench 1-100 $71.15 $33.33 4h16m 1h49m
Vibe Code Bench v1.1 $25.39 $7.89 1h07m 19m12s
CyberBench v1.1 $1.89 $3.31 16m57s 17m04s
Public Benefits Bench v1.1 $1.29 $1.20 21m09s 19m31s
Model details Claude Sonnet 5 Model details GPT-5.6 Terra