<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Vals AI Blog</title>
    <link>https://www.vals.ai/blogs</link>
    <atom:link href="https://www.vals.ai/blogs/rss.xml" rel="self" type="application/rss+xml" />
    <description>Research, benchmark releases, and analysis from Vals AI.</description>
    <language>en-us</language>
    <item>
      <title>Vals Web Search Index: Evaluating Search for Real Work</title>
      <link>https://www.vals.ai/blogs/web-search-index</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/web-search-index</guid>
      <pubDate>Fri, 02 Oct 2026 12:00:00 GMT</pubDate>
      <description>Why we built the Vals Web Search Index to measure whether a search tool helps an agent resolve real professional tasks.</description>
    </item>
    <item>
      <title>Has the Bitter Lesson Come for AI Detectors?</title>
      <link>https://www.vals.ai/blogs/ai-detection-benchmark</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/ai-detection-benchmark</guid>
      <pubDate>Wed, 30 Sep 2026 12:00:00 GMT</pubDate>
      <description>Benchmarking AI Detection Adversarially on Private Writing Samples</description>
    </item>
    <item>
      <title>A Lean Proof of the Thomson Problem for Seven Electrons</title>
      <link>https://www.vals.ai/blogs/thomson-n7-lean-proof</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/thomson-n7-lean-proof</guid>
      <pubDate>Mon, 28 Sep 2026 12:00:00 GMT</pubDate>
      <description>We asked ten Claude Sonnet 5.5 agents to prove in Lean that the pentagonal bipyramid is the lowest-energy way to place seven electrons on a sphere. In about 15 hours they produced a 17,895-line proof that Lean accepts, and a second kernel implementation confirms it.</description>
    </item>
    <item>
      <title>Evaluating AI Safety in Teen Conversations</title>
      <link>https://www.vals.ai/blogs/evaluating-ai-safety-in-teen-conversations</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/evaluating-ai-safety-in-teen-conversations</guid>
      <pubDate>Wed, 23 Sep 2026 12:00:00 GMT</pubDate>
      <description>A study of 648 simulated conversations examines how chatbots respond as teens disclose more, resist advice, and ask for help.</description>
    </item>
    <item>
      <title>A Faster Shortest Path Algorithm</title>
      <link>https://www.vals.ai/blogs/faster-shortest-path-algorithm</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/faster-shortest-path-algorithm</guid>
      <pubDate>Sun, 20 Sep 2026 12:00:00 GMT</pubDate>
      <description>We asked ten Claude Opus 5.5 agents to find a faster shortest-path algorithm and prove it in Lean. Within 15 hours, they produced C-HD: a formally verified improvement over the published bounds.</description>
    </item>
    <item>
      <title>AI Cheating is on the Rise</title>
      <link>https://www.vals.ai/blogs/cheating-on-the-rise</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/cheating-on-the-rise</guid>
      <pubDate>Tue, 15 Sep 2026 12:00:00 GMT</pubDate>
      <description>An integrity audit across BioMysteryBench, Terminal-Bench 2.1, and SWE-bench Verified shows that cheating increasingly complicates evaluation.</description>
    </item>
    <item>
      <title>Vals Environmental Impacts Report</title>
      <link>https://www.vals.ai/blogs/vals-environmental-impacts-report</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/vals-environmental-impacts-report</guid>
      <pubDate>Thu, 03 Sep 2026 12:00:00 GMT</pubDate>
      <description>Auditing the carbon, water, and energy costs of LLMs</description>
    </item>
    <item>
      <title>Claude Fable 5.1 Solves the Cyphral Distich</title>
      <link>https://www.vals.ai/blogs/fable-solves-cyphral-distich</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/fable-solves-cyphral-distich</guid>
      <pubDate>Mon, 31 Aug 2026 12:00:00 GMT</pubDate>
      <description>We gave Claude Fable 5.1 an open task: solve an unsolved 370-year-old cipher. It solved it within a day.</description>
    </item>
    <item>
      <title>Series A: Always a Higher Peak</title>
      <link>https://www.vals.ai/blogs/series-a</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/series-a</guid>
      <pubDate>Thu, 13 Aug 2026 12:00:00 GMT</pubDate>
      <description>Vals AI announces a $40M Series A at a $400M valuation led by a16z, alongside new releases including Vals Smith, Frontier Risk Benchmarks, and Vals 2.0.</description>
    </item>
    <item>
      <title>Child AI Safety</title>
      <link>https://www.vals.ai/blogs/child-ai-safety</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/child-ai-safety</guid>
      <pubDate>Wed, 12 Aug 2026 12:00:00 GMT</pubDate>
      <description>We are developing independent, youth-specific evaluations with clinicians and academic researchers.</description>
    </item>
    <item>
      <title>Vals Public Sector Memo</title>
      <link>https://www.vals.ai/blogs/vals-public-sector-memo</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/vals-public-sector-memo</guid>
      <pubDate>Thu, 02 Jul 2026 12:00:00 GMT</pubDate>
      <description>We&apos;re bringing independent AI evaluation to government.</description>
    </item>
    <item>
      <title>Measuring Healthcare AI Where It Actually Works</title>
      <link>https://www.vals.ai/blogs/measuring-healthcare-ai-where-it-actually-works</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/measuring-healthcare-ai-where-it-actually-works</guid>
      <pubDate>Mon, 23 Feb 2026 12:00:00 GMT</pubDate>
      <description>AI Models Excel at Clinical Documentation But Struggle With Medical Billing</description>
    </item>
    <item>
      <title>Behind the Scenes of Vibe Code Bench</title>
      <link>https://www.vals.ai/blogs/behind-the-scenes-of-vibe-code-bench</link>
      <guid isPermaLink="true">https://www.vals.ai/blogs/behind-the-scenes-of-vibe-code-bench</guid>
      <pubDate>Wed, 26 Nov 2025 12:00:00 GMT</pubDate>
      <description>Last week, we released Vibe Code Bench. This is our most ambitious benchmark to-date. I hope it will provide a useful signal for researchers and vibe-coders alike.</description>
    </item>
  </channel>
</rss>
