Compare Models

Select models (max 5)
Gemini 3 Pro (11/25)Mistral Large 4
Benchmarks

Vals Index *

Gemini 3 Pro (11/25)
N/A
Mistral Large 4
0.00%± 1.11
(44/44)

Legal Research Bench *

Gemini 3 Pro (11/25)
N/A
Mistral Large 4
0.00%± 3.23
(74/74)

Finance Agent (v2) *

Gemini 3 Pro (11/25)
N/A
Mistral Large 4
0.00%± 0.58
(75/75)

Tax Agent Bench *

Gemini 3 Pro (11/25)
N/A
Mistral Large 4
0.00%± 3.23
(66/66)

MedCode *

Gemini 3 Pro (11/25)
0.00%± 2.07
(105/105)
Mistral Large 4
0.00%± 2.17
(105/105)

Code Migration *

Gemini 3 Pro (11/25)
N/A
Mistral Large 4
0.00%± 4.22
(74/74)

Terminal-Bench 4.0

Gemini 3 Pro (11/25)
N/A
Mistral Large 4
0.00%± 0.88
(44/44)

Vibe Code Bench v1.1 *

Gemini 3 Pro (11/25)
0.00%± 3.06
(109/109)
Mistral Large 4
0.00%± 3.55
(109/109)

Overall performance

Performance on the Vals Index, a GDP-weighted aggregation of tasks across finance, coding, and law

2026-10-06

Comparison by Industry

Model performance on different sections of the economy.

Legal
87.03%
23.78%
Finance
70.82%
58.06%
Healthcare
62.12%
60.56%
Math
N/A
10.00%
Science
N/A
39.60%
Academic
89.76%
N/A
Education
47.62%
N/A
Coding
59.04%
38.25%

Cost Analysis

Per task and per token model pricing.

Vals Index · USD
N/A
$13.78
Cost / TestVals Index
N/A
$13.78
Input Cost/ 1M Tokens
$2.00
$1.36
Input Cache Read/ 1M Tokens
$0.20
$0.14
Output Cost/ 1M Tokens
$12.00
$4.18

Model Metadata

Basic information about each model.

Model provider
GoogleGoogle
MistralMistral AI
LatencyN/A6350.41s
Cost (In/Out)$2 / $12$1.36 / $4.18
Context Window1M512k
Max Output Token66k256k
Input Modality