Independent Evaluation, Unbiased Benchmarks

Testing AI on Real-World Tasks

We benchmark the world's leading AI models on economically valuable tasks such as finance, software, and frontier risk like cybersecurity, recursive self improvement and mental health. We run all of our own evaluations and create many of our benchmarks in-house.

Oct 07, 2026
Showing best model from each labView Full Results

Latest Reports

Recent benchmark releases and model evaluations.

Oct 08, 2026

Introducing SAFE-Teen: evaluating AI chatbots on teen mental health

Failure rate by risk area, without a teen-aware promptEach dot is one tested configuration. Lower is better.
Urgent safety3% - 34%
Health advice4% - 59%
Relationship boundaries10% - 50%
0%20%40%60%

Hover or tap a dot to follow one configuration across all three areas.

Industry Leaderboard

Model performance on different sections of the economy.

Industry
Benchmark

Vibe Code Bench v1.1

Benchmark data unavailable

Benchmark data not found

Model Performance Over Time

Tracking how foundation models improve with each release

AccuracyTime
Vals Index•Oct 09, 2026
Viewing 17 lab performance frontiers.
100Accuracy806040200
Feb '26Apr '26May '26Jul '26Sep '26