{
  "id": 9403174,
  "title": "OpenAI launches MentalHealthBench to evaluate AI mental health responses",
  "url": "https://urgent.news/2026/09/23/openai-launches-mentalhealthbench-to-evaluate-ai-mental-health",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-23T19:38:45.000Z",
  "source": {
    "name": "Investing.com",
    "slug": "investing-com",
    "url": "https://www.investing.com/news/stock-market-news/openai-launches-mentalhealthbench-to-evaluate-ai-mental-health-responses-93CH-4913784"
  },
  "original_language": "en",
  "account": "OpenAI launched MentalHealthBench on Wednesday, a new open benchmark designed with over 80 licensed mental health professionals from 22 nations. The benchmark aims to assess AI's performance in realistic mental health conversations, focusing on crucial behaviors such as safety, context-seeking, preserving user agency, and providing actionable guidance. Scenarios in the benchmark involve adults, teenagers, caregivers, and clinicians across various languages and regions.\n\nMentalHealthBench evaluates conversations in three categories: non-acute everyday interactions (53.5%), high-acuity situations with serious mental health concerns (18.2%), and emergencies requiring immediate safety action (28.3%). Each scenario was examined by a minimum of three experts who provided detailed rubric criteria to judge model responses, with scores ranging from -10 to +10 depending on their clinical significance. OpenAI assessed multiple AI models on the benchmark, with GPT-6 Astra achieving a score of 57.3%, GPT-6 Sol scoring 53.9%, and Claude Opus 5.5 scoring 52.4%.\n\nThe evaluation utilized GPT-5.6 Sol as an automated grader to measure model responses against the expert-written criteria. The data presented 95% confidence intervals for all results. Additionally, OpenAI conducted a separate study involving 44 adults from 16 countries who had utilized AI for mental health support. User feedback highlighted practical next steps and tone as desirable qualities, while experts stressed the importance of gathering relevant context and interpreting ambiguous situations.\n\nDespite the results, the benchmark's final scoring criteria were determined through expert consensus. OpenAI made the benchmark available for free to researchers, enabling them to scrutinize the methods, run their own evaluations, and build upon the work. The company also announced enhancements to ChatGPT's responses in sensitive conversations, expanded crisis resources, the introduction of Trusted Contact to facilitate connections with trusted individuals, and the launch of ChatGPT for Teens with additional safety measures.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}