Independent Evaluation, Unbiased Benchmarks

Testing AI on Real-World Tasks

We benchmark the world's leading AI models on rigorous, domain-specific tasks in finance, law, software, healthcare, and more. We run all of our own evaluations and create many of our benchmarks in-house.

Vals AI Updates

Fresh updates from our testing queue

model
07/09/2026

OpenAI's GPT-5.6 Sol and Terra evaluated across our benchmark suite

OpenAI's GPT-5.6 Sol and Terra evaluated across our benchmark suite

View Details

Benchmarks

Accuracy

Rankings

72.63%

± 1.07
2/ 36

72.19%

± 0.98
2/ 27

52.92%

± 4.35
2/ 29

64.38%

± 0.94
31/ 122

88.14%

± 4.08
1/ 15

72.34%

± 2.37
1/ 22

53.76%

± 0.85
7/ 36

48.08%

± 3.47
1/ 21

2.50%

± 0.83
12/ 22

95.20%

± 1.07
2/ 122
Contact us
Or send us an email at contact@vals.ai

License type:

Proprietary (contact us to get access)
Industry Partner
Academic

Read our methodology.

Industry Leaderboard

Independent benchmarks for industry-specific AI performance.

Industry
Benchmark

Model Performance Over Time

Tracking how foundation models improve with each release