Compare Models

Select models (max 5)
Claude Opus 5.5Claude Sonnet 5.5Claude Fable 5.1Claude Opus 5GPT-6 Astra
Benchmarks

Vals Index *

Claude Opus 5.5
0.00%± 0.94
(66/66)
Claude Sonnet 5.5
0.00%± 0.96
(66/66)
Claude Fable 5.1
0.00%± 1.08
(66/66)
Claude Opus 5
0.00%± 0.98
(66/66)
GPT-6 Astra
0.00%± 1.09
(66/66)

Legal Research Bench *

Claude Opus 5.5
0.00%± 3.48
(69/69)
Claude Sonnet 5.5
0.00%± 3.47
(69/69)
Claude Fable 5.1
0.00%± 3.46
(69/69)
Claude Opus 5
0.00%± 3.46
(69/69)
GPT-6 Astra
0.00%± 3.40
(69/69)

Finance Agent (v2) *

Claude Opus 5.5
0.00%± 0.17
(70/70)
Claude Sonnet 5.5
0.00%± 0.67
(70/70)
Claude Fable 5.1
0.00%± 2.06
(70/70)
Claude Opus 5
0.00%± 0.08
(70/70)
GPT-6 Astra
0.00%± 2.08
(70/70)

Tax Agent Bench *

Claude Opus 5.5
0.00%± 3.15
(61/61)
Claude Sonnet 5.5
0.00%± 2.98
(61/61)
Claude Fable 5.1
0.00%± 2.83
(61/61)
Claude Opus 5
0.00%± 2.96
(61/61)
GPT-6 Astra
0.00%± 3.15
(61/61)

MedCode *

Claude Opus 5.5
0.00%± 2.27
(102/102)
Claude Sonnet 5.5
0.00%± 2.12
(102/102)
Claude Fable 5.1
0.00%± 2.17
(102/102)
Claude Opus 5
0.00%± 1.99
(102/102)
GPT-6 Astra
0.00%± 2.13
(102/102)

Terminal-Bench Science

Claude Opus 5.5
0.00%± 6.02
(33/33)
Claude Sonnet 5.5
0.00%± 5.86
(33/33)
Claude Fable 5.1
0.00%± 5.71
(33/33)
Claude Opus 5
0.00%± 5.05
(33/33)
GPT-6 Astra
0.00%± 5.71
(33/33)

Code Migration *

Claude Opus 5.5
0.00%± 4.33
(69/69)
Claude Sonnet 5.5
0.00%± 4.26
(69/69)
Claude Fable 5.1
0.00%± 4.81
(69/69)
Claude Opus 5
0.00%± 4.37
(69/69)
GPT-6 Astra
0.00%± 4.22
(69/69)

Terminal-Bench 4.0

Claude Opus 5.5
0.00%± 1.01
(37/37)
Claude Sonnet 5.5
0.00%± 1.51
(37/37)
Claude Fable 5.1
0.00%± 3.07
(37/37)
Claude Opus 5
0.00%± 2.62
(37/37)
GPT-6 Astra
0.00%± 3.07
(37/37)

Vibe Code Bench v1.1 *

Claude Opus 5.5
0.00%± 1.53
(104/104)
Claude Sonnet 5.5
0.00%± 1.26
(104/104)
Claude Fable 5.1
0.00%± 1.57
(104/104)
Claude Opus 5
0.00%± 3.00
(104/104)
GPT-6 Astra
0.00%± 2.17
(104/104)

Overall performance

Performance on the Vals Index, a GDP-weighted aggregation of tasks across finance, coding, and law

2026-09-27

Comparison by Industry

Model performance on different sections of the economy.

Legal
27.12%
25.50%
50.16%
49.64%
22.42%
Finance
68.34%
69.07%
71.99%
70.89%
62.86%
Healthcare
70.61%
72.01%
72.40%
77.28%
68.20%
Math
100.00%
100.00%
100.00%
99.00%
99.00%
Science
59.13%
56.26%
41.02%
46.50%
66.04%
Academic
N/A
N/A
92.15%
91.64%
N/A
Education
45.83%
51.77%
48.53%
49.43%
46.37%
Coding
60.41%
74.58%
59.03%
61.52%
57.92%
Beta
34.31%
44.87%
42.45%
26.35%
51.90%
Social Mobility
70.64%
67.19%
74.90%
76.93%
N/A

Cost Analysis

Per task and per token model pricing.

Vals Index · USD
$32.77
$20.80
$28.92
$18.81
$19.09
Cost / TestVals Index
$32.77
$20.80
$28.92
$18.81
$19.09
Input Cost/ 1M Tokens
$4.00
$2.00
$10.00
$5.00
$10.00
Input Cache Write/ 1M Tokens
$5.00
$2.50
$12.50
$6.25
$12.50
Input Cache Read/ 1M Tokens
$0.40
$0.20
$0.25
$0.50
$1.00
Output Cost/ 1M Tokens
$20.00
$10.00
$50.00
$25.00
$50.00

Model Metadata

Basic information about each model.

Model provider
AnthropicAnthropic
AnthropicAnthropic
AnthropicAnthropic
AnthropicAnthropic
OpenAIOpenAI
Latency4320.73s4207.12s4576.31s3348.57s1510.53s
Cost (In/Out)$4 / $20$2 / $10$10 / $50$5 / $25$10 / $50
Context Window1M1M1M1M1M
Max Output Token128k128k128k128k128k
Input Modality