Claude Opus 5 vs GPT-6 Astra: Benchmark Comparison
Claude Opus 5 has the higher score on 12 of 24 shared benchmarks; GPT-6 Astra leads on 10.
The largest observed score gap is 71.67 pts on Time Horizon Index: KSP , where GPT-6 Astra leads.
Reported ±1 standard-error ranges overlap on 10 of 22 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Claude Opus 5 | GPT-6 Astra | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 63.67% ±0.96 | 63.13% ±1.19 | 0.55 pts | Reported ±1 SE ranges overlap |
| Vals RSI Index | 33.02% | 27.07% | 5.95 pts | Uncertainty comparison unavailable |
| Harvey's Legal Agent Benchmark | 6.67% ±1.65 | 5.42% ±1.17 | 1.25 pts | Reported ±1 SE ranges overlap |
| Legal Research Bench | 55.29% ±3.46 | 39.42% ±3.40 | 15.86 pts | Reported ±1 SE ranges do not overlap |
| EMB | 73.56% ±2.24 | 71.70% ±2.48 | 1.86 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 58.63% ±0.08 | 53.54% ±2.08 | 5.09 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 75.06% ±2.96 | 63.34% ±3.15 | 11.73 pts | Reported ±1 SE ranges do not overlap |
| MedCode | 63.57% ±1.99 | 48.49% ±2.13 | 15.08 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 90.98% ±1.92 | 87.91% ±1.94 | 3.08 pts | Reported ±1 SE ranges overlap |
| ProofBench v1.1 | 99.00% ±1.00 | 99.00% ±1.00 | 0.00 pts | Reported ±1 SE ranges overlap |
| BioMysteryBench | 79.26% ±2.06 | 79.26% ±0.98 | 0.00 pts | Reported ±1 SE ranges overlap |
| MysteryMechanism | 37.39% ±3.25 | 53.15% ±3.36 | 15.77 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench Science | 27.14% ±5.35 | 62.86% ±5.82 | 35.71 pts | Reported ±1 SE ranges do not overlap |
| SAGE | 49.43% ±3.29 | 46.37% ±3.44 | 3.06 pts | Reported ±1 SE ranges overlap |
| Code Migration | 57.47% ±4.37 | 67.74% ±4.22 | 10.27 pts | Reported ±1 SE ranges do not overlap |
| IOI | 84.33% ±9.96 | 100.00% ±0.00 | 15.67 pts | Reported ±1 SE ranges do not overlap |
| ProgramBench | 3.00% ±1.21 | 5.50% ±1.62 | 2.50 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench 4.0 | 53.53% ±1.34 | 59.60% ±4.40 | 6.06 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench 1-100 | 28.53% ±4.33 | 27.64% ±4.08 | 0.89 pts | Reported ±1 SE ranges overlap |
| Vibe Code Bench v1.1 | 88.40% ±3.00 | 89.59% ±2.17 | 1.19 pts | Reported ±1 SE ranges overlap |
| CyberBench v1.1 | 65.36% ±5.55 | 41.07% ±2.56 | 24.28 pts | Reported ±1 SE ranges do not overlap |
| SRE Bench | 12.21% ±2.03 | 56.87% ±3.07 | 44.66 pts | Reported ±1 SE ranges do not overlap |
| CUA-bench | 9.00% | 19.17% | 10.17 pts | Uncertainty comparison unavailable |
| Time Horizon Index: KSP | 18.83% ±0.00 | 90.50% ±0.00 | 71.67 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Claude Opus 5 average | GPT-6 Astra average |
|---|---|---|
| Legal | 30.98% | 22.42% |
| Finance | 69.08% | 62.86% |
| Healthcare | 77.28% | 68.20% |
| Math | 99.00% | 99.00% |
| Science | 47.93% | 65.09% |
| Education | 49.43% | 46.37% |
| Coding | 52.55% | 58.35% |
| Cyber | 38.79% | 48.97% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Claude Opus 5 cost | GPT-6 Astra cost | Claude Opus 5 latency | GPT-6 Astra latency |
|---|---|---|---|---|
| Vals Index | $19.31 | $18.46 | 57m41s | 27m48s |
| Vals RSI Index | $418.69 | $759.51 | 90h00m | 90h00m |
| Harvey's Legal Agent Benchmark | $23.67 | $26.16 | 55m37s | 24m49s |
| Legal Research Bench | $6.76 | $10.58 | 31m01s | 24m51s |
| EMB | $6.07 | $5.83 | 25m59s | 13m24s |
| Finance Agent (v2) | $5.12 | $6.82 | 9m58s | 11m33s |
| Tax Agent Bench | $5.14 | $5.80 | 16m48s | 15m01s |
| MedCode | N/A | N/A | 24.50s | 102.86s |
| MedScribe | N/A | N/A | 76.56s | 2m34s |
| ProofBench v1.1 | $2.11 | $1.71 | 9m51s | 5m08s |
| BioMysteryBench | $3.29 | $1.72 | 36m36s | 4m46s |
| MysteryMechanism | $3.28 | $1.56 | 16m53s | 8m09s |
| Terminal-Bench Science | $32.54 | $20.80 | 2h12m | 1h49m |
| SAGE | N/A | N/A | 56.17s | 41.63s |
| Code Migration | $60.51 | $44.36 | 2h51m | 58m30s |
| IOI | $16.48 | $6.50 | 56m42s | 31m40s |
| ProgramBench | $60.29 | $11.57 | 2h55m | 27m02s |
| Terminal-Bench 4.0 | $18.60 | $9.58 | 1h07m | 35m42s |
| Vibe Code Bench 1-100 | $41.49 | $84.72 | 1h22m | 3h26m |
| Vibe Code Bench v1.1 | $33.88 | $38.51 | 1h28m | 39m04s |
| CyberBench v1.1 | $2.87 | $1.84 | 19m25s | 4m10s |
| SRE Bench | $23.63 | $10.24 | 1h21m | 17m51s |
| CUA-bench | $232.23 | $1821.37 | 25.29s | 32.09s |
| Time Horizon Index: KSP | $1348.77 | $4206.59 | N/A | N/A |
Results available only for Claude Opus 5
- LegalBench
- MortgageTax
- TaxEval v2
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- LiveCodeBench
- SkillsBench
- SWE-bench
- Public Benefits Bench v1.1
Results available only for GPT-6 Astra
None.