OpenAI launches MentalHealthBench to evaluate AI mental health responses
OpenAI launched MentalHealthBench on Wednesday, a new open benchmark designed with over 80 licensed mental health professionals from 22 nations. The benchmark aims to assess AI's performance in realistic mental health conversations, focusing on crucial behaviors such as safety, context-seeking, preserving user agency, and providing actionable guidance. Scenarios in the benchmark involve adults, teenagers, caregivers, and clinicians across various languages and regions.
MentalHealthBench evaluates conversations in three categories: non-acute everyday interactions (53.5%), high-acuity situations with serious mental health concerns (18.2%), and emergencies requiring immediate safety action (28.3%). Each scenario was examined by a minimum of three experts who provided detailed rubric criteria to judge model responses, with scores ranging from -10 to +10 depending on their clinical significance.
OpenAI assessed multiple AI models on the benchmark, with GPT-6 Astra achieving a score of 57.3%, GPT-6 Sol scoring 53.9%, and Claude Opus 5.5 scoring 52.4%.
The evaluation utilized GPT-5.6 Sol as an automated grader to measure model responses against the expert-written criteria. The data presented 95% confidence intervals for all results. Additionally, OpenAI conducted a separate study involving 44 adults from 16 countries who had utilized AI for mental health support. User feedback highlighted practical next steps and tone as desirable qualities, while experts stressed the importance of gathering relevant context and interpreting ambiguous situations.
Despite the results, the benchmark's final scoring criteria were determined through expert consensus. OpenAI made the benchmark available for free to researchers, enabling them to scrutinize the methods, run their own evaluations, and build upon the work. The company also announced enhancements to ChatGPT's responses in sensitive conversations, expanded crisis resources, the introduction of Trusted Contact to facilitate connections with trusted individuals, and the launch of ChatGPT for Teens with additional safety measures.
Written by urgent.news from Investing.com's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.