Has the Bitter Lesson Come for AI Detectors?
Benchmarking AI Detection Adversarially on Private Writing Samples


A Lean Proof of the Thomson Problem for Seven Electrons
We asked ten Claude Sonnet 5.5 agents to prove in Lean that the pentagonal bipyramid is the lowest-energy way to place seven electrons on a sphere. In about 15 hours they produced a 17,895-line proof that Lean accepts, and a second kernel implementation confirms it.

Evaluating AI Safety in Teen Conversations
A study of 648 simulated conversations examines how chatbots respond as teens disclose more, resist advice, and ask for help.



A Faster Shortest Path Algorithm
We asked ten Claude Opus 5.5 agents to find a faster shortest-path algorithm and prove it in Lean. Within 15 hours, they produced C-HD: a formally verified improvement over the published bounds.

AI Cheating is on the Rise
An integrity audit across BioMysteryBench, Terminal-Bench 2.1, and SWE-bench Verified shows that cheating increasingly complicates evaluation.


Vals Environmental Impacts Report
Auditing the carbon, water, and energy costs of LLMs




Claude Fable 5.1 Solves the Cyphral Distich
We gave Claude Fable 5.1 an open task: solve an unsolved 370-year-old cipher. It solved it within a day.


Series A: Always a Higher Peak


Child AI Safety
We are developing independent, youth-specific evaluations with clinicians and academic researchers.



Vals Public Sector Memo
We're bringing independent AI evaluation to government.

