Compare Models
Select models (max 5)
Claude Opus 5GPT-6 Sol
| Vals IndexGDP-weighted benchmark | 63.67%±0.96 | 57.54%±1.01 |
| Legal Research BenchAgentic US legal research | 55.29%±3.46 | 28.85%±3.15 |
| Finance Agent (v2)Core financial analyst tasks | 58.63%±0.08 | 49.05%±0.58 |
| Tax Agent BenchAgentic US corporate tax research | 75.06%±2.96 | 53.05%±3.29 |
| MedCodeMedical billing code support | 63.57%±1.99 | 47.07%±2.12 |
| Terminal-Bench ScienceExpert-authored scientific research workflows | 27.14%±5.35 | 30.00%±5.52 |
| Code MigrationRewriting programs in new languages | 57.47%±4.37 | 57.20%±4.21 |
| Terminal-Bench 4.0Frontier-difficulty terminal tasks | 53.53%±1.34 | 44.44%±3.54 |
| Vibe Code Bench v1.1Building web apps from scratch | 88.40%±3.00 | 87.82%±2.53 |
Benchmarks
Vals Index *
Claude Opus 5
0.00%± 0.96
(43/43)GPT-6 Sol
0.00%± 1.01
(43/43)Legal Research Bench *
Claude Opus 5
0.00%± 3.46
(73/73)GPT-6 Sol
0.00%± 3.15
(73/73)Finance Agent (v2) *
Claude Opus 5
0.00%± 0.08
(74/74)GPT-6 Sol
0.00%± 0.58
(74/74)Tax Agent Bench *
Claude Opus 5
0.00%± 2.96
(65/65)GPT-6 Sol
0.00%± 3.29
(65/65)MedCode *
Claude Opus 5
0.00%± 1.99
(104/104)GPT-6 Sol
0.00%± 2.12
(104/104)Terminal-Bench Science
Claude Opus 5
0.00%