Model comparison

Claude Sonnet 5.5 vs GPT-5.6 Terra: Benchmark Comparison

Claude Sonnet 5.5 has the higher score on 15 of 17 shared benchmarks; GPT-5.6 Terra leads on 2.

The largest observed score gap is 28.57 pts on Terminal-Bench Science , where Claude Sonnet 5.5 leads.

Reported ±1 standard-error ranges overlap on 4 of 17 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5.5 GPT-5.6 Terra Gap Reported uncertainty
Vals Index 69.22% ±0.96 59.59% ±1.33 9.63 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 2.92% ±1.36 0.83% ±0.83 2.08 pts Reported ±1 SE ranges overlap
Legal Research Bench 48.08% ±3.47 41.35% ±3.42 6.73 pts Reported ±1 SE ranges overlap
EMB 75.71% ±2.44 66.20% ±3.11 9.51 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 58.10% ±0.67 54.44% ±2.07 3.67 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 73.39% ±2.98 65.20% ±3.19 8.20 pts Reported ±1 SE ranges do not overlap
MedCode 52.92% ±2.12 43.41% ±2.17 9.51 pts Reported ±1 SE ranges do not overlap
MedScribe 91.10% ±1.96 82.87% ±1.95 8.23 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 100.00% ±0.00 74.00% ±4.41 26.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 38.57% ±5.86 10.00% ±3.61 28.57 pts Reported ±1 SE ranges do not overlap
SAGE 51.77% ±3.41 47.00% ±3.40 4.77 pts Reported ±1 SE ranges overlap
Code Migration 69.83% ±4.26 47.80% ±4.28 22.02 pts Reported ±1 SE ranges do not overlap
IOI 83.06% ±3.74 87.61% ±6.53 4.56 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 53.03% ±1.51 26.26% ±0.51 26.77 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 92.39% ±1.26 74.59% ±4.20 17.80 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 59.58% ±5.21 72.08% ±5.41 12.50 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 67.19% ±1.22 62.38% ±1.26 4.80 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Sonnet 5.5 average GPT-5.6 Terra average
Index 69.22% 59.59%
Legal 25.50% 21.09%
Finance 69.07% 61.95%
Healthcare 72.01% 63.14%
Math 100.00% 74.00%
Science 38.57% 10.00%
Education 51.77% 47.00%
Coding 74.58% 59.07%
Beta 59.58% 72.08%
Social Mobility 67.19% 62.38%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5.5 cost GPT-5.6 Terra cost Claude Sonnet 5.5 latency GPT-5.6 Terra latency
Vals Index $20.80 $5.18 1h10m 28m50s
Harvey's Legal Agent Benchmark $16.61 $3.20 57m44s 17m05s
Legal Research Bench $14.83 $7.91 1h33m 1h10m
EMB $8.79 $2.19 36m45s 13m33s
Finance Agent (v2) $6.55 $3.63 31m43s 24m10s
Tax Agent Bench $9.25 $4.83 59m21s 1h01m
MedCode N/A N/A 3m59s 18.41s
MedScribe N/A N/A 5m07s 35.53s
ProofBench v1.1 $0.56 $1.27 5m48s 12m57s
Terminal-Bench Science $31.98 $10.58 3h22m 2h47m
SAGE N/A N/A 113.42s 40.06s
Code Migration $75.83 $8.13 3h20m 48m34s
IOI $7.52 $8.66 45m51s 1h23m
Terminal-Bench 4.0 $19.33 $5.91 2h06m 1h27m
Vibe Code Bench v1.1 $31.25 $7.89 1h01m 19m12s
CyberBench v1.1 $4.42 $3.31 33m53s 17m04s
Public Benefits Bench v1.1 $4.25 $1.20 1h18m 21m54s

Results available only for Claude Sonnet 5.5

  • BioMysteryBench
  • MysteryMechanism
  • SRE Bench

Results available only for GPT-5.6 Terra

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • ProgramBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
Model details Claude Sonnet 5.5 Model details GPT-5.6 Terra