About Vals

The independent evaluator of artificial intelligence. We build the benchmarks and evaluation infrastructure that measure whether models can do the work of lawyers, bankers, engineers, and doctors.

The Measurement Gap

Rapid progress in AI has been driven by a process researchers call hill-climbing: repeatedly defining a measurable objective, identifying shortcomings, and improving against them. Every major advance in AI, from image classification to question-answering to software development, has been measured this way. One hill at a time.

The pace of model development is outpacing the community’s ability to construct new hills. Trillions have been invested in generating intelligence, but comparatively little in measuring it. As a result, models summit old benchmarks in months. Test sets released openly are absorbed into pre-training corpora, which quietly invalidates the results built on them. And new benchmarks are often produced or run by the same companies that build the models, reported alongside cherry-picked examples and evaluation regimens tuned to score well.

AI has quickly become a trillion-dollar market without the independent measurement institutions that other markets of this scale, like finance and healthcare, require. The industry is worse off because of it.

  • Labs lack a credible way to demonstrate continued model progress.
  • Enterprises increasingly see pressure to adopt and spend on AI without means to quantify ROI.
  • Governments must develop their own expertise to measure frontier capabilities, cyber risk, and the pace of global competition.

We started Vals AI to solve this problem, as the independent evaluator of artificial intelligence.

What We Build

We build benchmarks that measure the ability of models to do the work of lawyers, bankers, engineers, and doctors. This takes partnering with reference institutions in each field to build a taxonomy of representative tasks, and developing new methodologies to automatically score the quality of generated work product. These benchmarks, including Finance Agent, have rapidly become industry standards.

Scores are based on our privately held test sets to preserve the integrity and signal of our results. This prevents training to the test set, unlike in open-source benchmarks.

We have built, and even open-sourced, the infrastructure we rely on to run these evaluations reproducibly and at scale across labs. This includes Valkyrie, a distributed system to run agentic benchmarks, and our model library, a free standard API to call models with.

We also produce benchmarks we believe will be broadly beneficial to the public good. For example, Public Benefits Bench, built with the Center for Civic Futures and Code for America, measures whether AI can be trusted to answer SNAP questions for the 37 million families who rely on the program.

We are hiring. Come join us.

Get in touch

Whether you build AI, evaluate it, or want to help shape a benchmark in your field, we would love to hear from and work with you.

Or send us an email at