Compare Models
Select models (max 5)
GPT-6.1 SolClaude Sonnet 5.5
| Vals IndexGDP-weighted benchmark | 61.15%±1.01 | 67.04%±0.92 |
| Legal Research BenchAgentic US legal research | 38.46%±3.38 | 48.08%±3.47 |
| Finance Agent (v2)Core financial analyst tasks | 52.03%±0.31 | 58.10%±0.67 |
| Tax Agent BenchAgentic US corporate tax research | 62.31%±3.20 | 73.39%±2.98 |
| MedCodeMedical billing code support | 48.84%±2.12 | 52.92%±2.12 |
| Terminal-Bench ScienceExpert-authored scientific research workflows | N/A | 38.57%±5.86 |
| Code MigrationRewriting programs in new languages | 65.12%±4.36 | 69.83%±4.26 |
| Terminal-Bench 4.0Frontier-difficulty terminal tasks | 55.05%±1.82 | 64.14%±1.01 |
| Vibe Code Bench v1.1Building web apps from scratch | 88.93%±1.89 | 92.39%±1.26 |
Benchmarks
Vals Index *
GPT-6.1 Sol
0.00%± 1.01
(30/30)Claude Sonnet 5.5
0.00%± 0.92
(30/30)Legal Research Bench *
GPT-6.1 Sol
0.00%± 3.38
(71/71)Claude Sonnet 5.5
0.00%± 3.47
(71/71)Finance Agent (v2) *
GPT-6.1 Sol
0.00%± 0.31
(72/72)Claude Sonnet 5.5
0.00%± 0.67
(72/72)Tax Agent Bench *
GPT-6.1 Sol
0.00%± 3.20
(63/63)Claude Sonnet 5.5
0.00%± 2.98
(63/63)MedCode *
GPT-6.1 Sol
0.00%± 2.12
(103/103)Claude Sonnet 5.5
0.00%± 2.12
(103/103)Terminal-Bench Science
GPT-6.1 Sol
N/AClaude Sonnet 5.5
0.00%± 5.86
(33/33)Code Migration *
GPT-6.1 Sol
0.00%± 4.36
(70/70)Claude Sonnet 5.5
0.00%± 4.26
(70/70)Terminal-Bench 4.0
GPT-6.1 Sol
0.00%± 1.82
(31/31)Claude Sonnet 5.5
0.00%± 1.01
(31/31)Vibe Code Bench v1.1 *
GPT-6.1 Sol
0.00%± 1.89
(105/105)Claude Sonnet 5.5
0.00%± 1.26
(105/105)Overall performance
Performance on the Vals Index, a GDP-weighted aggregation of tasks across finance, coding, and law