The Public Standard for
Real-World AI Benchmarks

Generic benchmarks only go so far.
Vals AI evaluates models on the real tasks each industry relies on.

Vals AI runs 30 independent AI benchmarks across legal, finance, healthcare, coding, math, and more. Each leaderboard compares large language models (LLMs) on accuracy, cost, and latency, and each methodology page explains how the benchmark is run.

As of October 2, 2026, Gemini 4 Argon ranks first on the Vals Index with 68.90%, followed by Claude Sonnet 5.5 (67.04%) and Claude Opus 5.5 (66.97%). See the full Vals Index leaderboard.

Vals Index

Finance benchmarks

Healthcare benchmarks

Math benchmarks

AIME

Challenging national math exam given to top high-school students

View Details

MATH 500

Academic math benchmark on probability, algebra, and trigonometry

View Details

MGSM

A multilingual benchmark for mathematical questions.

View Details

Science benchmarks

Academic benchmarks

GPQA Diamond

Graduate-level Google-Proof Q&A benchmark evaluating models on questions that require deep reasoning.

View Details

MMLU Pro

Academic multiple-choice benchmark covering 14 subjects including STEM, humanities, and social sciences.

View Details

MMMU Pro

Multimodal Multi-task Benchmark

View Details

Education benchmarks

Voice benchmarks

Coding benchmarks

Terminal-Bench 2.1

State-of-the-art set of difficult terminal-based tasks

View Details

SWE-bench Verified

Solving production software engineering tasks

View Details

LiveCodeBench

Our Implementation of the LiveCodeBench benchmark

View Details

Terminal-Bench 2.0

State-of-the-art set of difficult terminal-based tasks

View Details

Cyber benchmarks

Games benchmarks

Social Mobility benchmarks