Claude Opus 5 vs GPT-6 Astra: Benchmark Comparison

Claude Opus 5 has the higher score on 12 of 24 shared benchmarks; GPT-6 Astra leads on 10.

The largest observed score gap is 71.67 pts on Time Horizon Index: KSP , where GPT-6 Astra leads.

Reported ±1 standard-error ranges overlap on 10 of 22 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Opus 5 GPT-6 Astra Gap Reported uncertainty
Vals Index 63.67% ±0.96 63.13% ±1.19 0.55 pts Reported ±1 SE ranges overlap
Vals RSI Index 33.02% 27.07% 5.95 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 6.67% ±1.65 5.42% ±1.17 1.25 pts Reported ±1 SE ranges overlap
Legal Research Bench 55.29% ±3.46 39.42% ±3.40 15.86 pts Reported ±1 SE ranges do not overlap
EMB 73.56% ±2.24 71.70% ±2.48 1.86 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 58.63% ±0.08 53.54% ±2.08 5.09 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 75.06% ±2.96 63.34% ±3.15 11.73 pts Reported ±1 SE ranges do not overlap
MedCode 63.57% ±1.99 48.49% ±2.13 15.08 pts Reported ±1 SE ranges do not overlap
MedScribe 90.98% ±1.92 87.91% ±1.94 3.08 pts Reported ±1 SE ranges overlap
ProofBench v1.1 99.00% ±1.00 99.00% ±1.00 0.00 pts Reported ±1 SE ranges overlap
BioMysteryBench 79.26% ±2.06 79.26% ±0.98 0.00 pts Reported ±1 SE ranges overlap
MysteryMechanism 37.39% ±3.25 53.15% ±3.36 15.77 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 27.14% ±5.35 62.86% ±5.82 35.71 pts Reported ±1 SE ranges do not overlap
SAGE 49.43% ±3.29 46.37% ±3.44 3.06 pts Reported ±1 SE ranges overlap
Code Migration 57.47% ±4.37 67.74% ±4.22 10.27 pts Reported ±1 SE ranges do not overlap
IOI 84.33% ±9.96 100.00% ±0.00 15.67 pts Reported ±1 SE ranges do not overlap
ProgramBench 3.00% ±1.21 5.50% ±1.62 2.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 53.53% ±1.34 59.60% ±4.40 6.06 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 28.53% ±4.33 27.64% ±4.08 0.89 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 88.40% ±3.00 89.59% ±2.17 1.19 pts Reported ±1 SE ranges overlap
CyberBench v1.1 65.36% ±5.55 41.07% ±2.56 24.28 pts Reported ±1 SE ranges do not overlap
SRE Bench 12.21% ±2.03 56.87% ±3.07 44.66 pts Reported ±1 SE ranges do not overlap
CUA-bench 9.00% 19.17% 10.17 pts Uncertainty comparison unavailable
Time Horizon Index: KSP 18.83% ±0.00 90.50% ±0.00 71.67 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Opus 5 average GPT-6 Astra average
Legal 30.98% 22.42%
Finance 69.08% 62.86%
Healthcare 77.28% 68.20%
Math 99.00% 99.00%
Science 47.93% 65.09%
Education 49.43% 46.37%
Coding 52.55% 58.35%
Cyber 38.79% 48.97%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Opus 5 cost GPT-6 Astra cost Claude Opus 5 latency GPT-6 Astra latency
Vals Index $19.31 $18.46 57m41s 27m48s
Vals RSI Index $418.69 $759.51 90h00m 90h00m
Harvey's Legal Agent Benchmark $23.67 $26.16 55m37s 24m49s
Legal Research Bench $6.76 $10.58 31m01s 24m51s
EMB $6.07 $5.83 25m59s 13m24s
Finance Agent (v2) $5.12 $6.82 9m58s 11m33s
Tax Agent Bench $5.14 $5.80 16m48s 15m01s
MedCode N/A N/A 24.50s 102.86s
MedScribe N/A N/A 76.56s 2m34s
ProofBench v1.1 $2.11 $1.71 9m51s 5m08s
BioMysteryBench $3.29 $1.72 36m36s 4m46s
MysteryMechanism $3.28 $1.56 16m53s 8m09s
Terminal-Bench Science $32.54 $20.80 2h12m 1h49m
SAGE N/A N/A 56.17s 41.63s
Code Migration $60.51 $44.36 2h51m 58m30s
IOI $16.48 $6.50 56m42s 31m40s
ProgramBench $60.29 $11.57 2h55m 27m02s
Terminal-Bench 4.0 $18.60 $9.58 1h07m 35m42s
Vibe Code Bench 1-100 $41.49 $84.72 1h22m 3h26m
Vibe Code Bench v1.1 $33.88 $38.51 1h28m 39m04s
CyberBench v1.1 $2.87 $1.84 19m25s 4m10s
SRE Bench $23.63 $10.24 1h21m 17m51s
CUA-bench $232.23 $1821.37 25.29s 32.09s
Time Horizon Index: KSP $1348.77 $4206.59 N/A N/A

Results available only for Claude Opus 5

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Public Benefits Bench v1.1

Results available only for GPT-6 Astra

None.

Model details Claude Opus 5 Model details GPT-6 Astra