Compare Models

Select models (max 5)
o3 MiniMistral Large 4
Benchmarks

Vals Index *

o3 Mini
N/A
Mistral Large 4
0.00%± 1.11
(44/44)

Legal Research Bench *

o3 Mini
N/A
Mistral Large 4
0.00%± 3.23
(74/74)

Finance Agent (v2) *

o3 Mini
N/A
Mistral Large 4
0.00%± 0.58
(75/75)

Tax Agent Bench *

o3 Mini
N/A
Mistral Large 4
0.00%± 3.23
(66/66)

MedCode *

o3 Mini
N/A
Mistral Large 4
0.00%± 2.17
(105/105)

Code Migration *

o3 Mini
N/A
Mistral Large 4
0.00%± 4.22
(74/74)

Terminal-Bench 4.0

o3 Mini
N/A
Mistral Large 4
0.00%± 0.88
(44/44)

Vibe Code Bench v1.1 *

o3 Mini
N/A
Mistral Large 4
0.00%± 3.55
(109/109)

Overall performance

Performance on the Vals Index, a GDP-weighted aggregation of tasks across finance, coding, and law

2026-10-06

Comparison by Industry

Model performance on different sections of the economy.

Legal
71.54%
23.78%
Finance
69.42%
58.06%
Healthcare
N/A
60.56%
Math
N/A
10.00%
Science
N/A
39.60%
Academic
77.10%
N/A
Coding
71.48%
38.25%

Cost Analysis

Per task and per token model pricing.

Vals Index · USD
N/A
$13.78
Cost / TestVals Index
N/A
$13.78
Input Cost/ 1M Tokens
$1.10
$1.36
Input Cache Read/ 1M Tokens
$0.55
$0.14
Output Cost/ 1M Tokens
$4.40
$4.18

Model Metadata

Basic information about each model.

Model provider
OpenAIOpenAI
MistralMistral AI
LatencyN/A6350.41s
Cost (In/Out)$1.1 / $4.4$1.36 / $4.18
Context Window200k512k
Max Output Token100k256k
Input Modality