Introducing MentalHealthBench
OpenAI released MentalHealthBench, an open benchmark built with 80+ licensed mental health experts to evaluate AI responses across mental health conversations.
OpenAI introduced MentalHealthBench, an open benchmark measuring how AI systems respond in realistic mental health conversations spanning non-acute, high-acuity, and emergency scenarios. It was co-created with more than 80 licensed psychologists and psychiatrists from 22 countries speaking 19 languages, who wrote rubric criteria weighted from -10 to +10 for each synthetic conversation. An automated grader, GPT-5.6 Sol, scores model responses against the expert criteria, with each conversation reviewed by at least three experts. OpenAI reports steady improvement across frontier models, positioning the benchmark as progress toward safer well-being support.