OpenAI has introduced MentalHealthBench, a new open benchmark designed to evaluate how AI systems respond in realistic mental health conversations. Developed in collaboration with a global cohort of more than 80 licensed psychologists and psychiatrists from 22 countries, the benchmark aims to bridge the gap in understanding how models perform across a wide range of mental health-related interactions, beyond just emergency scenarios.
Product LaunchesOpenAIMentalHealthBench
OpenAI Introduces MentalHealthBench to Evaluate AI in Mental Health Conversations
The benchmark consists of 1,215 synthetic conversations that cover various levels of acuity, including non-acute, high-acuity, and emergency situations. It also incorporates diverse user profiles, such as adults, teenagers, caregivers, and clinicians, across multiple languages and cultural contexts. To ensure clinical validity, the dataset utilizes privacy-preserving techniques to reflect real-world usage patterns of ChatGPT.
MentalHealthBench evaluates model capabilities through ten behavioral axes, including safety, seeking context, preserving user agency, and providing actionable guidance. Each conversation is paired with a rubric created by experts, where criteria are assigned weights to reward beneficial behaviors and penalize harmful ones. Grading is performed using an automated LLM-based evaluator, GPT-5.6 Sol, at high reasoning effort.
The research conducted alongside the benchmark release suggests that while frontier models show steady improvement in navigating mental health situations, significant challenges remain. Specifically, models still face difficulties in adequately seeking context and calibrating the level of urgency required for different situations. OpenAI has released the benchmark openly to allow researchers and developers to examine their methods and improve future models.
Sources
- Introducing MentalHealthBench (OpenAI News, 2026-09-23)
- MentalHealthBenchの論文