Proprietary

Vals Index

Updated 8/13/2026

A single measure of AI's potential economic impact — agentic model performance across finance, coding, and legal tasks, weighted by each sector's share of U.S. GDP.

Vals IndexGDP-weighted benchmark
ACCURACY

Motivation

As AI capabilities rapidly advance, understanding their potential to transform economic sectors has become critical for organizations making deployment decisions. Unlike existing aggregated metrics that treat all capabilities equally, the Vals Index is designed to reflect the potential impact of AI models on the U.S. economy. We accomplish this by computing a weighted average of model performance across key sectors, where each sector’s weight is proportional to its share of U.S. GDP.

Vals AI has developed a comprehensive suite of benchmarks measuring AI models’ ability to perform real-world tasks across finance, coding, and legal work — benchmarks that no other party has full access to. The Vals Index aggregates five private and two public benchmarks to provide high signal into the real-world tradeoffs between capability, latency, and cost of deploying AI systems.


Results

Industry Average Accuracy Comparison

Key Takeaways

AI models are advancing rapidly in their ability to handle complex, real-world tasks across critical economic sectors. The results demonstrate that frontier models are becoming increasingly capable at automating work in finance, coding, and legal services — domains that collectively represent a substantial portion of economic activity.

Claude Opus 5 leads the Vals Index at 67.21%, ahead of Claude Fable 5 at 66.04% and GPT-5.6 Sol at 63.71%. GPT-5.6 Luna reaches 59.88% at $0.62 per test, 3.8 points behind the third-place model at less than a twentieth of its cost.


Methodology

Benchmark Selection, Economic Weighting, and Formula

The Vals Index aggregates performance across three sectors — Finance, Coding, and Legal — each weighted in proportion to its share of U.S. GDP, using Bureau of Economic Analysis value-added-by-industry data (as published on FRED). While this represents a vast oversimplification of how AI might impact the economy, it provides a useful proxy for measuring the potential economic significance of model capabilities:

Finance: Finance & Insurance, ~8.0% of U.S. GDP

Coding: Information sector (the closest GDP proxy for software work), ~5.6% of U.S. GDP

Legal: Legal services, ~1.2% of U.S. GDP

  • Legal Research Bench: Case and statute research with citation-backed answers
  • HLAB: Harvey’s Legal Agent Benchmark, long-horizon legal work product creation

Each sector is the average of its benchmarks, and the sectors are combined with the following formula:

Finance = AVG(FinanceAgentV2, EMB)
Coding = AVG(TBench, VibeCodeBench, CodeMigration)
Legal = AVG(LegalResearch, HLAB)
Vals_Index = (8.0 * Finance + 5.6 * Coding + 1.2 * Legal) / 14.8

Component Scores

Each benchmark column on the index is that benchmark’s own published standalone score — the same number, under the same methodology. The exception is Code Migration: there, the index scores a fixed subset of the published run: 50 of the 120 CLI migration tasks, plus all 10 COBOL tasks, weighted 75% CLI and 25% COBOL to match the standalone benchmark’s balance. This subset was chosen by Monte Carlo search, scored against the full 120-task leaderboard. The chosen subset reproduces the full 120-task result to Spearman 0.99 and 0.8 points of mean absolute accuracy error, with no model moving more than three ranks; on models held out of the selection, mean absolute error is 1.1 points.


Updates

8/13/2026

Released Vals Index v2:

  • Added EMB in place of CorpFin. A private benchmark for building complex financial models in Excel.
  • Added Code Migration to the coding bucket. A private benchmark involving porting projects to a set of four languages, including COBOL modernization.
  • Added Legal Research Bench to the returning Legal sector. A private benchmark involving answering legal questions grounded in case law.
  • Added HLAB to the returning Legal sector. Harvey’s Legal Agent Benchmark, long-horizon legal work product creation.
  • Removed SWE-Bench Verified from the coding bucket. SWE-Bench has become saturated. Coding now averages Terminal-Bench 2.1, Vibe Code Bench, and Code Migration equally.

5/27/2026

  • Updated the coding bucket to Terminal-Bench 2.1. The Vals Index now evaluates coding performance with Terminal-Bench 2.1 while keeping the same 0.25 coding weight.

5/13/2026

  • Swapped Finance Agent to Finance Agent v2. Finance now uses the Finance Agent v2 index subset, averaging three runs per model.

5/4/2026

  • Added Vibe Code Bench to the coding bucket. Coding is now a weighted average of three benchmarks (SWE-Bench Verified 0.25, Terminal-Bench 2.0 0.25, Vibe Code Bench 0.5), giving end-to-end app-building tasks half of the coding signal. VCB is evaluated on a 22-task subset selected for coverage across UI, data, and workflow patterns.
  • Removed the Law sector (CaseLaw). The CaseLaw benchmark had become saturated, and was no longer providing useful differentiation between models. Consequently, it was removed from the index. The denominator was rebalanced from 3.7 → 3.4 to reflect the dropped 0.3 law weight, and the Industry Average chart no longer displays a Law column.