Compare Models

Select models (max 5)
Claude Haiku 5.5Mistral Large 4
Benchmarks

Vals Index *

Claude Haiku 5.5
0.00%± 1.41
(45/45)
Mistral Large 4
0.00%± 1.11
(45/45)

Legal Research Bench *

Claude Haiku 5.5
0.00%± 3.44
(75/75)
Mistral Large 4
0.00%± 3.23
(75/75)

Finance Agent (v2) *

Claude Haiku 5.5
0.00%± 2.13
(76/76)
Mistral Large 4
0.00%± 0.58
(76/76)

Tax Agent Bench *

Claude Haiku 5.5
0.00%± 3.22
(67/67)
Mistral Large 4
0.00%± 3.23
(67/67)

MedCode *

Claude Haiku 5.5
N/A
Mistral Large 4
0.00%± 2.17
(105/105)

Code Migration *

Claude Haiku 5.5
N/A
Mistral Large 4
0.00%± 4.22
(74/74)

Terminal-Bench 4.0

Claude Haiku 5.5
0.00%± 2.20
(45/45)
Mistral Large 4
0.00%± 0.88
(45/45)

Vibe Code Bench v1.1 *

Claude Haiku 5.5
0.00%± 1.73
(110/110)
Mistral Large 4
0.00%± 3.55
(110/110)

Overall performance

Performance on the Vals Index, a GDP-weighted aggregation of tasks across finance, coding, and law

2026-10-07

Comparison by Industry

Model performance on different sections of the economy.

Legal
22.26%
23.78%
Finance
58.45%
58.06%
Healthcare
N/A
60.56%
Math
86.00%
10.00%
Science
N/A
39.60%
Coding
57.69%
38.25%

Cost Analysis

Per task and per token model pricing.

Vals Index · USD
$2.991
$13.78
Cost / TestVals Index
$2.991
$13.78
Input Cost/ 1M Tokens
$0.10
$1.36
Input Cache Write/ 1M Tokens
$0.125
N/A
Input Cache Read/ 1M Tokens
$0.01
$0.14
Output Cost/ 1M Tokens
$0.50
$4.18

Model Metadata

Basic information about each model.

Model provider
AnthropicAnthropic
MistralMistral AI
Latency2836.53s6350.41s
Cost (In/Out)$0.1 / $0.5$1.36 / $4.18
Context Window1M512k
Max Output Token128k256k
Input Modality