Proprietary

Vals Multimodal Index

Updated 5/28/2026

Benchmark consisting of a weighted performance across finance, coding, and education tasks. Showing the potential impact that LLMs can have on the economy.

As of May 28, 2026, Claude Opus 4.8 ranks first on the Vals Multimodal Index v1.1 with 70.71%, followed by GPT 5.5 (67.77%) and Claude Opus 4.7 (67.36%).

Vals Multimodal Index
ACCURACY

Vals Multimodal Index leaderboard

Rank Model Accuracy Cost / Test Input / Output Cost Duration
1 Claude Opus 4.8 70.71% $1.95 $5 / $25 15m14s
2 GPT 5.5 67.77% $3.47 $5 / $30 12m49s
3 Claude Opus 4.7 67.36% $4.33 $5 / $25 9m02s
4 Gemini 3.5 Flash 62.29% $1.06 $1.5 / $9 4m32s
5 Claude Sonnet 4.6 60.78% $1.57 $3 / $15 8m43s
6 Kimi K2.6 56.79% $0.52 $0.95 / $4 14m41s
7 Gemini 3.1 Pro Preview (02/26) 55.75% $0.97 $2 / $12 5m42s
8 GPT 5.4 Mini 53.30% $0.56 $0.75 / $4.5 16m23s
9 Gemini 3 Flash (12/25) 51.98% $0.29 $0.5 / $3 4m32s
10 Qwen 3.6 Plus 50.74% $1.03 $0.5 / $3 10m50s
11 GPT 5.4 Nano 47.48% $0.26 $0.2 / $1.25 15m24s
12 Grok 4.3 43.44% $0.52 $1.25 / $2.5 9m11s
13 Claude Haiku 4.5 (Thinking) 42.35% $0.40 $1 / $5 4m44s
14 Gemini 3.1 Flash Lite Preview 40.47% $0.12 $0.25 / $1.5 3m28s
15 Grok 4.20 (Reasoning) 38.70% $0.36 $2 / $6 2m30s

Results

Industry Average Accuracy Comparison
Multimodal Vals Index

Methodology

This archived version uses Terminal-Bench 2.0 in the coding bucket, before Terminal-Bench 2.1 replaced it in v1.2.

Finance (weight: 2.0): ~$2T contribution to U.S. GDP

Coding (weight: 1.4): ~$1.4T contribution to U.S. GDP

Education (weight: 0.3): ~$270B contribution to U.S. GDP

  • SAGE: Grading handwritten student work in mathematics
Coding = 0.25 * SWE_Bench + 0.25 * TBench + 0.5 * VibeCodeBench
Vals_Multimodal_Index = (2.0 * AVG(CorpFin, FinanceAgent, MortgageTax) + 1.4 * Coding + 0.3 * SAGE) / 3.7