Claude Opus 4.8 vs GPT-6 Astra: Benchmark Comparison
Claude Opus 4.8 has the higher score on 6 of 16 shared benchmarks; GPT-6 Astra leads on 10.
The largest observed score gap is 78.67 pts on Time Horizon Index: KSP , where GPT-6 Astra leads.
Reported ±1 standard-error ranges overlap on 5 of 15 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Claude Opus 4.8 | GPT-6 Astra | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 55.10% ±1.01 | 63.13% ±1.19 | 8.02 pts | Reported ±1 SE ranges do not overlap |
| Vals RSI Index | 23.98% | 27.07% | 3.09 pts | Uncertainty comparison unavailable |
| Harvey's Legal Agent Benchmark | 9.58% ±2.15 | 5.42% ±1.17 | 4.17 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 43.75% ±3.45 | 39.42% ±3.40 | 4.33 pts | Reported ±1 SE ranges overlap |
| EMB | 69.37% ±2.61 | 71.70% ±2.48 | 2.33 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 53.92% ±0.16 | 53.54% ±2.08 | 0.38 pts | Reported ±1 SE ranges overlap |
| Tax Agent Bench | 64.91% ±3.17 | 63.34% ±3.15 | 1.57 pts | Reported ±1 SE ranges overlap |
| MedCode | 53.22% ±2.17 | 48.49% ±2.13 | 4.73 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 85.75% ±1.93 | 87.91% ±1.94 | 2.15 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench Science | 4.29% ±2.44 | 62.86% ±5.82 | 58.57 pts | Reported ±1 SE ranges do not overlap |
| SAGE | 54.79% ±3.35 | 46.37% ±3.44 | 8.42 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 47.25% ±4.18 | 67.74% ±4.22 | 20.49 pts | Reported ±1 SE ranges do not overlap |
| ProgramBench | 1.00% ±0.70 | 5.50% ±1.62 | 4.50 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench 4.0 | 23.23% ±1.34 | 59.60% ±4.40 | 36.36 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 82.72% ±3.08 | 89.59% ±2.17 | 6.87 pts | Reported ±1 SE ranges do not overlap |
| Time Horizon Index: KSP | 11.83% ±0.00 | 90.50% ±0.00 | 78.67 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Claude Opus 4.8 average | GPT-6 Astra average |
|---|---|---|
| Legal | 26.67% | 22.42% |
| Finance | 62.73% | 62.86% |
| Healthcare | 69.49% | 68.20% |
| Science | 4.29% | 62.86% |
| Education | 54.79% | 46.37% |
| Coding | 38.55% | 55.61% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Claude Opus 4.8 cost | GPT-6 Astra cost | Claude Opus 4.8 latency | GPT-6 Astra latency |
|---|---|---|---|---|
| Vals Index | $13.14 | $18.46 | 39m48s | 27m48s |
| Vals RSI Index | $465.25 | $759.51 | 90h00m | 90h00m |
| Harvey's Legal Agent Benchmark | $10.22 | $26.16 | 24m08s | 24m49s |
| Legal Research Bench | $2.82 | $10.58 | 11m54s | 24m51s |
| EMB | $12.06 | $5.83 | 43m52s | 13m24s |
| Finance Agent (v2) | $4.22 | $6.82 | 9m07s | 11m33s |
| Tax Agent Bench | $1.69 | $5.80 | 7m59s | 15m01s |
| MedCode | N/A | N/A | 105.92s | 102.86s |
| MedScribe | N/A | N/A | 82.36s | 2m34s |
| Terminal-Bench Science | $23.14 | $20.80 | 1h55m | 1h49m |
| SAGE | N/A | N/A | 3m06s | 41.63s |
| Code Migration | $30.51 | $44.36 | 1h15m | 58m30s |
| ProgramBench | $31.27 | $11.57 | 1h32m | 27m02s |
| Terminal-Bench 4.0 | $17.14 | $9.58 | 1h11m | 35m42s |
| Vibe Code Bench v1.1 | $26.88 | $38.51 | 1h16m | 39m04s |
| Time Horizon Index: KSP | $1499.32 | $4206.59 | N/A | N/A |
Results available only for Claude Opus 4.8
- LegalBench
- MortgageTax
- TaxEval v2
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- LiveCodeBench
- SkillsBench
- SWE-bench
- Public Benefits Bench v1.1
Results available only for GPT-6 Astra
- ProofBench v1.1
- BioMysteryBench
- MysteryMechanism
- IOI
- Vibe Code Bench 1-100
- CyberBench v1.1
- SRE Bench
- CUA-bench