Compare Models

Select models (max 5)
Muse SparkMistral Large 4
Benchmarks

Vals Index *

Muse Spark
N/A
Mistral Large 4
0.00%± 1.11
(44/44)

Legal Research Bench *

Muse Spark
N/A
Mistral Large 4
0.00%± 3.23
(74/74)

Finance Agent (v2) *

Muse Spark
N/A
Mistral Large 4
0.00%± 0.58
(75/75)

Tax Agent Bench *

Muse Spark
N/A
Mistral Large 4
0.00%± 3.23
(66/66)

MedCode *

Muse Spark
0.00%± 2.24
(105/105)
Mistral Large 4
0.00%± 2.17
(105/105)

Code Migration *

Muse Spark
N/A
Mistral Large 4
0.00%± 4.22
(74/74)

Terminal-Bench 4.0

Muse Spark
N/A
Mistral Large 4
0.00%± 0.88
(44/44)

Vibe Code Bench v1.1 *

Muse Spark
0.00%± 3.97
(109/109)
Mistral Large 4
0.00%± 3.55
(109/109)

Overall performance

Performance on the Vals Index, a GDP-weighted aggregation of tasks across finance, coding, and law

2026-10-06

Comparison by Industry

Model performance on different sections of the economy.

Legal
84.22%
23.78%
Finance
77.68%
58.06%
Healthcare
68.61%
60.56%
Math
N/A
10.00%
Science
N/A
39.60%
Academic
88.12%
N/A
Coding
47.04%
38.25%

Cost Analysis

Per task and per token model pricing.

Vals Index · USD
N/A
$13.78
Cost / TestVals Index
N/A
$13.78
Input Cost/ 1M Tokens
N/A
$1.36
Input Cache Read/ 1M Tokens
N/A
$0.14
Output Cost/ 1M Tokens
N/A
$4.18

Model Metadata

Basic information about each model.

Model provider
MetaMeta
MistralMistral AI
LatencyN/A6350.41s
Cost (In/Out)N/A$1.36 / $4.18
Context Window1M512k
Max Output Token131k256k
Input Modality