Archived Benchmark

This index scored the Vals Index v1.2 composition plus the multimodal SAGE and Mortgage Tax tasks. The Vals Index has since moved to its v2 composition, so we no longer run this variant on new model releases. Previous results are preserved here for posterity.

Proprietary

Vals Multimodal Index

Updated 8/11/2026

Benchmark consisting of a weighted performance across finance, coding, and education tasks. Showing the potential impact that LLMs can have on the economy.

As of August 11, 2026, Claude Fable 5 ranks first on the Vals Multimodal Index with 74.15%, followed by Claude Opus 5 (73.90%) and Kimi K3 (73.42%).

Vals Multimodal IndexGDP-weighted multimodal index
ACCURACY

Vals Multimodal Index leaderboard

Rank Model Accuracy Cost / Test Input / Output Cost Duration
1 Claude Fable 5 74.15% $7.97 $10 / $50 12m49s
2 Claude Opus 5 73.90% $6.14 $5 / $25 14m14s
3 Kimi K3 73.42% $1.70 $3 / $15 15m06s
4 GPT-5.6 Sol 72.64% $5.36 $4 / $20 9m11s
5 Claude Opus 4.8 70.89% $5.46 $5 / $25 16m20s
6 Muse Spark 1.2 69.80% $0.51 $1.25 / $4.25 7m26s
7 GPT-5.6 Luna 69.06% $0.16 $0.2 / $1.2 6m47s
8 Claude Sonnet 5 68.83% $6.53 $2 / $10 16m52s
9 GPT 5.5 68.07% $3.32 $5 / $30 8m24s
10 Claude Opus 4.7 67.36% $4.44 $5 / $25 9m24s
11 Muse Spark 1.1 66.74% $0.36 $1.25 / $4.25 4m48s
12 Qwen 3.8 Max 65.39% $1.93 $2 / $6 33m16s
13 Gemini 3.6 Flash 65.08% $1.02 $1.5 / $7.5 8m56s
14 GPT-5.6 Terra 65.07% $0.80 $2 / $12 4m18s
15 Grok 4.5 63.42% $0.91 $2 / $6 5m19s
16 Gemini 3.5 Flash 62.93% $1.01 $1.5 / $9 4m26s
17 Claude Sonnet 4.6 60.57% $1.58 $3 / $15 8m56s
18 MiniMax-M3 59.97% $1.08 $0.6 / $2.4 17m56s
19 Kimi K2.6 56.43% $0.52 $0.95 / $4 15m04s
20 Gemini 3.1 Pro Preview (02/26) 56.07% $0.98 $2 / $12 5m23s
21 GPT 5.4 Mini 54.22% $0.49 $0.75 / $4.5 13m08s
22 Qwen 3.7 Plus 53.89% $0.26 $0.4 / $1.6 10m29s
23 MiMo V2.5 52.77% $0.03 $0.14 / $0.28 7m10s
24 Gemini 3 Flash (12/25) 52.19% $0.29 $0.5 / $3 4m20s
25 Qwen 3.6 Plus 51.52% $1.12 $0.5 / $3 10m43s
26 Inkling Small 50.05% $0.14 $0.3 / $1.2 11m49s
27 Inkling 49.40% $0.63 $1 / $4.05 15m29s
28 GPT 5.4 Nano 47.64% $0.23 $0.2 / $1.25 11m11s
29 Grok 4.3 43.29% $0.50 $1.25 / $2.5 5m39s
30 Claude Haiku 4.5 (Thinking) 42.88% $0.41 $1 / $5 4m43s
31 Gemini 3.1 Flash Lite Preview 41.35% $0.12 $0.25 / $1.5 2m32s
32 Grok 4.20 (Reasoning) 39.06% $0.40 $2 / $6 2m51s
33 Mistral Medium 3.5 34.77% $3.45 $1.5 / $7.5 12m15s

Motivation

As AI capabilities rapidly advance, understanding their potential to transform economic sectors has become increasingly critical for organizations making deployment decisions. Unlike existing aggregated metrics that treat all capabilities equally, the Vals Index is designed to reflect the potential economic impact of AI models on the U.S. economy. We accomplish this by computing a weighted average of model performance across key sectors, where the weights correspond to each sector’s contribution to the U.S. economy in trillions of dollars.

Vals AI has developed a comprehensive suite of benchmarks measuring AI models’ ability to perform real-world tasks across finance, software engineering, and education. These benchmarks were designed to evaluate practical performance on actual professional workflows, making them well-suited for assessing economic impact. The Vals Multimodal Index leverages this existing work to provide a high-signal measure that accounts for the real-world tradeoffs between capability, latency, and cost that practitioners face when deploying AI systems.


Results

Industry Average Accuracy Comparison

Key Takeaways

Kimi K3 is the only open-weight model in the leading tier of this composite of hard, image-inclusive economic tasks. Claude Fable 5 leads at 74.15%, with Claude Opus 5 at 73.90% and Kimi K3 at 73.42%, all within a point of one another; GPT-5.6 Sol follows at 72.64%.


Methodology

Benchmark Selection, Economic Weighting, and Formula

The Vals Index aggregates performance across three major economic sectors, weighted by their approximate contribution to U.S. GDP. Market size estimations were computed based on data from the Federal Reserve Economic Data (FRED) and the Bureau of Labor Statistics. While this represents a vast oversimplification of how AI might impact the economy, it provides a useful proxy for measuring the potential economic significance of model capabilities:

Finance (weight: 2.0): ~$2T contribution to U.S. GDP

Coding (weight: 1.4): ~$1.4T contribution to U.S. GDP

Education (weight: 0.3): ~$270B contribution to U.S. GDP

  • SAGE: Grading handwritten student work in mathematics

These weights combine in the following formula:

Coding = 0.25 * SWE_Bench + 0.25 * TBench + 0.5 * VibeCodeBench
Vals_Multimodal_Index = (2.0 * AVG(CorpFin, FinanceAgent, MortgageTax) + 1.4 * Coding + 0.3 * SAGE) / 3.7

The denominator (3.7) normalizes the index to a 0-100 scale, where the score represents the weighted average performance across sectors proportional to their economic contribution.

Subset Selection Process

To enable efficient and cost-effective evaluation while maintaining strong correlation with full benchmark performance, we developed representative subsets for three benchmarks:

Selection Methodology: To balance evaluation efficiency with accuracy, we created representative subsets for select benchmarks using a sampling process that maximizes correlation with full benchmark scores. We validated this approach using holdout models to ensure that subset performance reliably predicts full benchmark results.

Benchmark-Specific Subsets:

  • SWE-bench Verified: 33 randomly sampled instances from each difficulty level (categorized by solution time: <15min, 15min-1hr, 1-4hr, >4hr), plus all 3 instances from the hardest category
  • CorpFin: 3 randomly selected questions per unique document from the original test set
  • Finance Agent v2: 13-model multimodal index subset evaluated with three runs per model
  • Vibe Code Bench: 22 representative app-building tasks selected to cover a range of UI, data, and workflow patterns

Full Benchmarks:

This methodology ensures the Multimodal Vals Index provides a rapid, cost-effective evaluation framework while maintaining the predictive validity needed for reliable model comparison.


Updates

5/27/2026

  • Updated the coding bucket to Terminal-Bench 2.1. The Vals Multimodal Index now evaluates coding performance with Terminal-Bench 2.1 while keeping the same 0.25 coding weight.