Blog Announcements

Child AI Safety

We are developing independent, youth-specific evaluations with clinicians and academic researchers.

Andrea MockGlenn Parham
Andrea Mock & Glenn Parham 08/12/2026
Child AI Safety

Today, we’re sharing a new area of work at Vals focused on AI and mental health, starting with children and teenagers. As more young people turn to AI for help, we think this is an area that needs substantially better evaluation.

RAND recently found that nearly 1 in 5 Americans ages 12 to 21 had used an AI chatbot for mental health advice, up more than 40% in a year. Common Sense Media, looking more broadly at AI companions, found that 72% of teens have used one and 52% use them at least a few times a month.

As usage grows, these systems are also facing greater scrutiny over potential real-world harms. For example, the family of 14-year-old Sewell Setzer III alleged that a Character AI chatbot fostered an intense emotional dependency and contributed to his death by suicide in 2024. Character AI and Google later agreed to settle the lawsuit, and families in several other states have brought related cases alleging serious harm to minors.

There are already signs that these conversations are difficult to handle safely. In one JMIR Mental Health study, 10 therapy and companion chatbots endorsed dangerous or unwise proposals from fictional teenagers in 19 of 60 cases. Separate tests by Common Sense Media found that ChatGPT, Claude, Gemini, and Meta AI performed better on explicit single-turn safety prompts than in longer teen mental-health conversations, where safeguards broke down, and models missed warning signs.

Major AI companies are also changing their safeguards and evaluation practices. OpenAI has introduced teen-specific safeguards and dynamic multi-turn evaluations for mental health, emotional reliance, and self-harm. Character AI removed open-ended chat for users under 18. Anthropic restricts Claude.ai to adults, imposes additional requirements on developers serving minors, and tests Claude in extended multi-turn self-harm scenarios.

Policymakers are also increasingly asking how AI systems used by minors can be independently assessed. California’s SB 243 requires new safety protocols and reporting for companion chatbots, while federal efforts include the KIDS Act and the SAFE KIDS Act, which would require ongoing risk assessments and annual independent child-safety audits.

As evaluations play a larger role in determining whether these systems are safe, the quality of the evaluations themselves matters.

What we are doing at Vals AI

At Vals, we are developing independent, youth-specific evaluations with clinicians and academic researchers. The work currently focuses on three areas:

  • Self-harm and crisis situations
  • Models acting like doctors, therapists, or other inappropriate authorities
  • Unhealthy emotional attachment and dependence on AI

We want to measure the trajectory of an interaction, including cases where risk is indirect, a user resists advice, warning signs build over time, or individually reasonable responses add up to an unhealthy exchange. We aim to go beyond single-turn refusals and ground these evaluations in real-world conversations.

What we’re seeing so far:

  • Recognizing a crisis is easier than responding usefully. In one comparison of various models, all models recognized an explicit suicide disclosure, but their follow-up differed substantially with some offered only a vague referral, while others also assessed immediate danger, and helped the adolescent reach a real person.
  • Important failures appeared in ordinary-seeming questions. Some conversations began as requests for drug education, symptom interpretation, or help finding resources. The risk emerged gradually as the model answered questions about diagnosis, reporting, or treatment with more confidence and detail than it should have.
  • The way a teenager frames the relationship matters. Models often maintained a boundary when asked directly to be a therapist or romantic partner, but were less consistent when the same role developed through coaching, grief support, fictional characters, or requests to become the teenager’s only source of support.

In high-stakes settings, AI safety may depend on interactions rather than individual outputs. We are starting with children and teenagers because the stakes are unusually high and appropriate behavior is highly context-sensitive. But the same evaluation problem applies more broadly whenever people turn to general-purpose AI for emotional support.

We have more detailed findings from our evaluations across foundation models, which we’re sharing first with policymakers given the sensitivity of this work. We’re also looking to collaborate with clinicians and foundation model labs on future evaluations and plan to publish more on the methodology and results soon. If you’re interested in collaborating, get in touch or join our mailing list to stay updated.