Model comparison

Claude Sonnet 5.5 vs GPT-6 Astra: Benchmark Comparison

Claude Sonnet 5.5 has the higher score on 13 of 19 shared benchmarks; GPT-6 Astra leads on 6.

The largest observed score gap is 27.14 pts on Terminal-Bench Science , where GPT-6 Astra leads.

Reported ±1 standard-error ranges overlap on 9 of 19 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5.5 GPT-6 Astra Gap Reported uncertainty
Vals Index 69.22% ±0.96 66.61% ±1.09 2.61 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 2.92% ±1.36 5.42% ±1.17 2.50 pts Reported ±1 SE ranges overlap
Legal Research Bench 48.08% ±3.47 39.42% ±3.40 8.65 pts Reported ±1 SE ranges do not overlap
EMB 75.71% ±2.44 71.70% ±2.48 4.01 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 58.10% ±0.67 53.54% ±2.08 4.56 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 73.39% ±2.98 63.34% ±3.15 10.06 pts Reported ±1 SE ranges do not overlap
MedCode 52.92% ±2.12 48.49% ±2.13 4.43 pts Reported ±1 SE ranges do not overlap
MedScribe 91.10% ±1.96 87.91% ±1.94 3.19 pts Reported ±1 SE ranges overlap
ProofBench v1.1 100.00% ±0.00 99.00% ±1.00 1.00 pts Reported ±1 SE ranges overlap
BioMysteryBench 81.11% ±0.64 79.26% ±0.98 1.85 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 49.10% ±3.36 53.15% ±3.36 4.05 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 38.57% ±5.86 65.71% ±5.71 27.14 pts Reported ±1 SE ranges do not overlap
SAGE 51.77% ±3.41 46.37% ±3.44 5.41 pts Reported ±1 SE ranges overlap
Code Migration 69.83% ±4.26 67.74% ±4.22 2.09 pts Reported ±1 SE ranges overlap
IOI 83.06% ±3.74 100.00% ±0.00 16.94 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 53.03% ±1.51 57.07% ±3.07 4.04 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 92.39% ±1.26 89.59% ±2.17 2.80 pts Reported ±1 SE ranges overlap
CyberBench v1.1 59.58% ±5.21 41.07% ±2.56 18.51 pts Reported ±1 SE ranges do not overlap
SRE Bench 30.15% ±2.84 56.87% ±3.07 26.72 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Sonnet 5.5 average GPT-6 Astra average
Index 69.22% 66.61%
Legal 25.50% 22.42%
Finance 69.07% 62.86%
Healthcare 72.01% 68.20%
Math 100.00% 99.00%
Science 56.26% 66.04%
Education 51.77% 46.37%
Coding 74.58% 78.60%
Beta 44.87% 48.97%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5.5 cost GPT-6 Astra cost Claude Sonnet 5.5 latency GPT-6 Astra latency
Vals Index $20.80 $19.09 1h10m 25m11s
Harvey's Legal Agent Benchmark $16.61 $26.16 57m44s 24m49s
Legal Research Bench $14.83 $10.58 1h33m 24m51s
EMB $8.79 $5.83 36m45s 13m24s
Finance Agent (v2) $6.55 $6.82 31m43s 11m33s
Tax Agent Bench $9.25 $5.80 59m21s 15m01s
MedCode N/A N/A 3m59s 102.86s
MedScribe N/A N/A 5m07s 2m34s
ProofBench v1.1 $0.56 $1.71 5m48s 5m08s
BioMysteryBench $3.74 $1.72 18m06s 4m46s
MysteryMechanism $3.45 $1.56 23m30s 8m09s
Terminal-Bench Science $31.98 $15.80 3h22m 1h30m
SAGE N/A N/A 113.42s 41.63s
Code Migration $75.83 $44.36 3h20m 58m30s
IOI $7.52 $6.50 45m51s 31m40s
Terminal-Bench 4.0 $19.33 $8.21 2h06m 42m02s
Vibe Code Bench v1.1 $31.25 $38.51 1h01m 39m04s
CyberBench v1.1 $4.42 $1.84 33m53s 4m10s
SRE Bench $26.52 $13.50 1h55m 17m51s

Results available only for Claude Sonnet 5.5

  • Public Benefits Bench v1.1

Results available only for GPT-6 Astra

  • Vals RSI Index
  • ProgramBench
  • Vibe Code Bench 1-100
  • CUA-bench
  • Time Horizon Index: KSP
Model details Claude Sonnet 5.5 Model details GPT-6 Astra