Rapid progress in AI has been driven by a process known to researchers as “hill-climbing.” This entails repeatedly defining a measurable objective, identifying shortcomings, and improving against them. Every major advance in AI, from image classification to question-answering and software development, has been measured this way. One hill at a time.
What we predicted when we started Vals, and now see coming into reality, is that the pace of model development is outpacing the community’s ability to construct new hills. Trillions have been invested in generating intelligence, but comparatively little in measuring it. As a result, models summit old benchmarks in months and tests get leaked into the training corpora. New benchmarks are often produced or run by the same companies that build the models.
AI has quickly become a trillion-dollar market without the independent measurement institutions that other markets of this scale, like finance and healthcare, require. The industry is worse off because of it.
Labs lack a credible way to demonstrate continued model progress. Enterprises increasingly see pressure to adopt and spend on AI without the means to quantify ROI. Governments must develop their own expertise to measure frontier capabilities, cyber risk, and the pace of global competition.
We started Vals AI to solve this problem as the independent evaluator of artificial intelligence.
We build benchmarks that measure the ability of models to do the work of lawyers, bankers, engineers, and doctors. This takes partnering with reference institutions in each field to build a taxonomy of representative tasks, and new methodologies to automatically score the quality of generated work product. We’ve built and even open sourced the infra we rely on to run these evaluations reproducibly and at scale across labs. Scores are based on our privately held test sets to preserve the integrity and signal of our results.
This approach is quickly becoming the standard. Our results have been cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI. Some of the biggest enterprise AI deployments use Vals to choose which models to build on and measure their products against the frontier. We have supported the Department of Commerce and members of Congress working on AI policy. Our revenue has grown 8x compared with all of 2025, our customer base has doubled, and our incredible team has tripled in six months.
Today we’re announcing our $40M Series A at a $400M valuation, led by a16z, with participation from existing investors 8VC and BloombergBeta and new investors HRT Ventures and Next Ladder Ventures.
Alongside the raise, we’re proud to share three new releases:
- Vals Smith. Now generally available, Vals Smith lets anyone create a custom coding benchmark from any GitHub repo, with 120 free credits to get started.
- Frontier Risk Benchmarks. We’re releasing our RSI Index in collaboration with CoreWeave, a new cyber benchmark built with leading academics, and announcing our initial work in mental health.
- Vals 2.0. We’ve completely rebuilt the Vals AI website and launched a new Vals Index with more coverage of the economy.
Build your first benchmark with Vals Smith or, if you’re building your nth, consider joining us.
There is always a higher peak.