Proprietary

Vals Multimodal Index

Updated 5/4/2026

Benchmark consisting of a weighted performance across finance, coding, and education tasks. Showing the potential impact that LLMs can have on the economy.

As of May 4, 2026, Claude Opus 4.7 ranks first on the Vals Multimodal Index v1 with 71.10%, followed by GPT 5.5 (69.84%) and GPT 5.4 (xhigh) at 65.17%.

Vals Multimodal Index
ACCURACY

Vals Multimodal Index leaderboard

Rank Model Accuracy Cost / Test Input / Output Cost Duration
1 Claude Opus 4.7 71.10% $4.00 $5 / $25 8m55s
2 GPT 5.5 69.84% $3.46 $5 / $30 13m30s
3 GPT 5.4 (xhigh) 65.17% $2.02 $2.5 / $15 18m08s
4 Claude Opus 4.6 (Thinking) 65.16% $1.84 $5 / $25 7m39s
5 Claude Sonnet 4.6 64.00% $1.43 $3 / $15 8m03s
6 GPT 5.2 63.75% $3.03 $1.75 / $14 17m24s
7 Kimi K2.6 59.10% $0.44 $0.95 / $4 12m19s
8 Gemini 3.1 Pro Preview (02/26) 58.62% $0.91 $2 / $12 5m54s
9 GPT 5.4 Mini 55.93% $0.49 $0.75 / $4.5 14m02s
10 Claude Opus 4.5 (Thinking) 55.22% $5.36 $5 / $25 8m35s
11 Gemini 3 Flash (12/25) 54.33% $0.26 $0.5 / $3 5m59s
12 Qwen 3.6 Plus 53.98% $0.96 $0.5 / $3 10m10s
13 Gemini 3 Pro (11/25) 53.08% $1.22 $2 / $12 27m29s
14 GPT 5.1 52.70% $0.63 $1.25 / $10 9m34s
15 Kimi K2.5 52.01% $0.21 $0.6 / $3 10m13s
16 Claude Sonnet 4.5 (Thinking) 51.98% $1.41 $3 / $15 10m54s
17 GPT 5.4 Nano 50.80% $0.24 $0.2 / $1.25 15m08s
18 Qwen 3.6 27B 50.18% $0.93 $0.6 / $3.6 32m11s
19 GPT 5 49.93% $0.42 $1.25 / $10 10m31s
20 Qwen 3.5 Plus 48.79% $0.77 $0.4 / $2.4 14m39s
21 Grok 4.3 47.04% $0.46 $1.25 / $2.5 9m32s
22 Claude Haiku 4.5 (Thinking) 46.58% $0.37 $1 / $5 4m34s
23 GPT 5 Mini 46.19% $0.09 $0.25 / $2 6m31s
24 Gemini 3.1 Flash Lite Preview 43.86% $0.11 $0.25 / $1.5 3m38s
25 Grok 4.20 (Reasoning) 42.97% $0.30 $2 / $6 110.25s
26 Gemini 2.5 Pro 42.41% $0.36 $1.25 / $10 12m05s
27 Grok 4 Fast (Reasoning) 38.37% $0.05 $0.2 / $0.5 5m04s
28 Grok 4.1 Fast (Reasoning) 38.18% $0.05 $0.2 / $0.5 3m19s

Results

Industry Average Accuracy Comparison
Multimodal Vals Index

Methodology

This archived version uses the original Vals Multimodal Index finance component, before the Finance Agent v2 index subset replaced it in v1.1.

Finance (weight: 2.0): ~$2T contribution to U.S. GDP

Coding (weight: 1.4): ~$1.4T contribution to U.S. GDP

Education (weight: 0.3): ~$270B contribution to U.S. GDP

  • SAGE: Grading handwritten student work in mathematics
Coding = 0.25 * SWE_Bench + 0.25 * TBench + 0.5 * VibeCodeBench
Vals_Multimodal_Index = (2.0 * AVG(CorpFin, FinanceAgent, MortgageTax) + 1.4 * Coding + 0.3 * SAGE) / 3.7