Independent Evaluation, Unbiased Benchmarks

Testing AI on Real-World Tasks

We benchmark the world's leading AI models on rigorous, domain-specific tasks in finance, law, software, healthcare, and more. We run all of our own evaluations and create many of our benchmarks in-house.

Vals AI Updates

Fresh updates from our testing queue

benchmark
07/23/2026

Introducing the Web Search Index: native provider search vs. independent web-search tools

Introducing the Web Search Index: native provider search vs. independent web-search tools

View Details

System

Accuracy

48.45%

± 2.03

46.94%

± 2.10

45.24%

± 2.01

43.58%

± 1.98

41.36%

± 1.93

38.75%

± 1.95

37.04%

± 1.85

35.54%

± 1.93
Showing top 10 models from the benchmark. Visit the benchmark page to view more

Industry Leaderboard

Independent benchmarks for industry-specific AI performance.

Industry
Benchmark

Model Performance Over Time

Tracking how foundation models improve with each release

85%73%61%49%37%25%
Oct '25Nov '25Jan '26Mar '26Apr '26Jun '26Aug '26