Claude Opus 5.5 vs GPT-6 Astra: Benchmark Comparison

Claude Opus 5.5 has the higher score on 15 of 24 shared benchmarks; GPT-6 Astra leads on 8.

The largest observed score gap is 23.28 pts on SRE Bench , where GPT-6 Astra leads.

Reported ±1 standard-error ranges overlap on 12 of 22 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Opus 5.5 GPT-6 Astra Gap Reported uncertainty
Vals Index 66.97% ±0.89 63.13% ±1.19 3.85 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 37.31% 27.07% 10.24 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 3.75% ±1.47 5.42% ±1.17 1.67 pts Reported ±1 SE ranges overlap
Legal Research Bench 50.48% ±3.48 39.42% ±3.40 11.06 pts Reported ±1 SE ranges do not overlap
EMB 75.94% ±2.38 71.70% ±2.48 4.24 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 58.59% ±0.17 53.54% ±2.08 5.05 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 70.50% ±3.15 63.34% ±3.15 7.16 pts Reported ±1 SE ranges do not overlap
MedCode 49.80% ±2.27 48.49% ±2.13 1.31 pts Reported ±1 SE ranges overlap
MedScribe 91.43% ±1.93 87.91% ±1.94 3.52 pts Reported ±1 SE ranges overlap
ProofBench v1.1 100.00% ±0.00 99.00% ±1.00 1.00 pts Reported ±1 SE ranges overlap
BioMysteryBench 79.26% ±0.98 79.26% ±0.98 0.00 pts Reported ±1 SE ranges overlap
MysteryMechanism 49.55% ±3.36 53.15% ±3.36 3.60 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 47.14% ±6.01 62.86% ±5.82 15.71 pts Reported ±1 SE ranges do not overlap
SAGE 45.83% ±3.36 46.37% ±3.44 0.54 pts Reported ±1 SE ranges overlap
Code Migration 66.65% ±4.33 67.74% ±4.22 1.09 pts Reported ±1 SE ranges overlap
IOI 95.06% ±4.94 100.00% ±0.00 4.94 pts Reported ±1 SE ranges overlap
ProgramBench 18.50% ±2.75 5.50% ±1.62 13.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 65.15% ±0.00 59.60% ±4.40 5.56 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 30.36% ±4.83 27.64% ±4.08 2.72 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 90.29% ±1.53 89.59% ±2.17 0.70 pts Reported ±1 SE ranges overlap
CyberBench v1.1 55.36% ±5.13 41.07% ±2.56 14.28 pts Reported ±1 SE ranges do not overlap
SRE Bench 33.59% ±2.92 56.87% ±3.07 23.28 pts Reported ±1 SE ranges do not overlap
CUA-bench 14.00% 19.17% 5.17 pts Uncertainty comparison unavailable
Time Horizon Index: KSP 91.33% ±0.00 90.50% ±0.00 0.83 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Opus 5.5 average GPT-6 Astra average
Legal 27.12% 22.42%
Finance 68.34% 62.86%
Healthcare 70.61% 68.20%
Math 100.00% 99.00%
Science 58.65% 65.09%
Education 45.83% 46.37%
Coding 61.00% 58.35%
Cyber 44.47% 48.97%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Opus 5.5 cost GPT-6 Astra cost Claude Opus 5.5 latency GPT-6 Astra latency
Vals Index $32.14 $18.46 1h19m 27m48s
Vals RSI Index $411.86 $759.51 90h00m 90h00m
Harvey's Legal Agent Benchmark $21.38 $26.16 49m00s 24m49s
Legal Research Bench $25.16 $10.58 1h52m 24m51s
EMB $11.02 $5.83 35m06s 13m24s
Finance Agent (v2) $9.22 $6.82 33m08s 11m33s
Tax Agent Bench $15.09 $5.80 1h07m 15m01s
MedCode N/A N/A 4m07s 102.86s
MedScribe N/A N/A 7m22s 2m34s
ProofBench v1.1 $0.96 $1.71 5m51s 5m08s
BioMysteryBench $3.64 $1.72 12m19s 4m46s
MysteryMechanism $4.62 $1.56 19m38s 8m09s
Terminal-Bench Science $19.12 $20.80 2h28m 1h49m
SAGE N/A N/A 114.20s 41.63s
Code Migration $112.97 $44.36 3h19m 58m30s
IOI $5.26 $6.50 17m15s 31m40s
ProgramBench $68.05 $11.57 2h24m 27m02s
Terminal-Bench 4.0 $13.20 $9.58 1h04m 35m42s
Vibe Code Bench 1-100 $44.12 $84.72 2h28m 3h26m
Vibe Code Bench v1.1 $57.92 $38.51 1h36m 39m04s
CyberBench v1.1 $3.48 $1.84 18m11s 4m10s
SRE Bench $34.65 $10.24 1h37m 17m51s
CUA-bench $188.04 $1821.37 14.71s 32.09s
Time Horizon Index: KSP $1303.37 $4206.59 N/A N/A

Results available only for Claude Opus 5.5

  • Public Benefits Bench v1.1

Results available only for GPT-6 Astra

None.

Model details Claude Opus 5.5 Model details GPT-6 Astra