Independent Evaluation, Unbiased Benchmarks

Testing AI on Real-World Tasks

We benchmark the world's leading AI models on rigorous, domain-specific tasks in finance, law, software, healthcare, and more. We run all of our own evaluations and create many of our benchmarks in-house.

Vals AI Updates

Fresh updates from our testing queue

model
07/23/2026

Anthropic's Claude Opus 5 evaluated across our benchmark suite

Anthropic's Claude Opus 5 evaluated across our benchmark suite

View Details
Show scores with fallbacks counted as failures

Benchmarks

Accuracy

Rankings

74.82%

± 1.35
2/ 40

73.90%

± 1.24
2/ 29

57.47%

± 4.37
1/ 33

73.19%

± 0.87
1/ 126

40.68%

± 2.54
18/ 20

73.56%

± 2.24
2/ 28

58.63%

± 0.08
1/ 40

55.29%

± 3.46
1/ 27

6.67%

± 1.65
7/ 27

93.43%

± 1.24
4/ 126
Contact us
Or send us an email at contact@vals.ai

License type:

Proprietary (contact us to get access)
Industry Partner
Academic

Read our methodology.

Industry Leaderboard

Independent benchmarks for industry-specific AI performance.

Industry
Benchmark

Model Performance Over Time

Tracking how foundation models improve with each release

85%73%61%49%37%25%
Oct '25Nov '25Jan '26Mar '26Apr '26Jun '26Aug '26