The Public Standard for
Real World AI Performance
Generic benchmarks only go so far.
Vals AI evaluates models on the
real tasks each industry relies on.
Vals Index
Vals Multimodal Index
Benchmark consisting of a weighted performance across finance, coding, and education tasks. Showing the potential impact that LLMs can have on the economy.
Legal
CaseLaw v2
Private question-answer benchmark over Canadian court-cases.
Finance
CorpFin v2
A private benchmark evaluating understanding of long-context credit agreements
MortgageTax
Evaluating reading and understanding tax certificates as images
TaxEval v2
A Vals-created set of questions and responses to tax questions
Healthcare
MedQA
Evaluating language model bias in medical questions.
Math
AIME
Challenging national math exam given to top high-school students
MATH 500
Academic math benchmark on probability, algebra, and trigonometry
MGSM
A multilingual benchmark for mathematical questions.
Science
Academic
GPQA Diamond
Graduate-level Google-Proof Q&A benchmark evaluating models on questions that require deep reasoning.
MMLU Pro
Academic multiple-choice benchmark covering 14 subjects including STEM, humanities, and social sciences.
MMMU Pro
Multimodal Multi-task Benchmark
Education
Voice
Coding
SWE-bench Verified
Solving production software engineering tasks
LiveCodeBench
Our Implementation of the LiveCodeBench benchmark
Terminal-Bench 2.0
State-of-the-art set of difficult terminal-based tasks
Social Mobility
Public Benefits Bench v1.1
Can AI help people navigate SNAP benefits?
Top Models
Claude Opus 5
Claude Fable 5.1
Claude Fable 5
Public Benefits Bench v1
Can AI help people navigate SNAP benefits?