Independent Evaluation, Unbiased Benchmarks

Testing AI on Real-World Tasks

We benchmark the world's leading AI models on economically valuable tasks such as finance, software, and frontier risk like cybersecurity, recursive self improvement and mental health. We run all of our own evaluations and create many of our benchmarks in-house.

Latest Reports

Recent benchmark releases and model evaluations.

Aug 13, 2026

Introducing the RSI Index: can a model do the research that builds the next model?

Industry Leaderboard

Model performance on different sections of the economy.

Industry
Benchmark

Vibe Code Bench v1.1

Benchmark data unavailable

Benchmark data not found

Model Performance Over Time

Tracking how foundation models improve with each release

AccuracyTime
Vals IndexAug 13, 2026
100Accuracy806040200
Oct '25Dec '25Feb '26May '26Jul '26