Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations

Open benchmark of 1,215 synthetic mental health conversations, each paired with weighted rubric criteria written and adjudicated by a cohort of more than 80 licensed psychiatrists and psychologists from 22 countries. Conversations span non-acute, high-acuity and emergency situations, four user profiles (adults, teens aged 13 to 17, clinicians, caregivers) and eleven languages, and responses are graded by an LLM judge against the expert rubric and decomposed into ten behavioural axes. The paper reports scores for 17 models from OpenAI, Anthropic, Google, xAI and Meta and compares expert rubrics with rubrics written by 44 adult users of AI for emotional support.

Publisher

OpenAI

Published

23 Sept 2026

Added

today

DOI

Key Findings

  • Task-clipped scores: GPT-6 Astra 57.3, GPT-6 Sol 53.9, Claude Opus 5.5 52.4, GPT-6 Luna 50.2, Muse Spark 1.3 48.6, GPT-5.6 Sol 47.0, Claude Fable 5.1 46.4, Claude Sonnet 5 44.5, Grok 4.7 41.3, Gemini 3.8 Flash 35.5, GPT-4o (March 2025) 32.1, Gemini 3.1 Pro 32.1, Gemini 2.5 Pro 29.5.
  • Composition: 53.5% non-acute, 18.2% high-acuity, 28.3% emergency conversations; 68.1% adult, 21.2% teen, 5.8% clinician, 4.9% caregiver profiles; 159 suicide and self-harm and 115 psychosis conversations; 312 non-English conversations led by Spanish (105) and Hindi (54).
  • GPT-4o (March 2025) scores 36.5 on non-acute but 24.1 on emergency conversations, while GPT-6 Astra holds 56.6, 57.9 and 58.3 across the three acuity levels.
  • On the teen subset (age stated via a system message) Claude Opus 5.5 scores highest at 57.0, ahead of GPT-6 Astra (55.9); Grok 4.7 scores 40.5 and Gemini 2.5 Pro 29.1.
  • Rubrics written by 44 adult users of AI for emotional support align with expert rubrics on 25.7% of weight; 39.1% of weight is expert-only (mostly guardrails), 34.2% user-only (actionable next steps and tone) and 1.0% directly contradictory.
  • Context-seeking (66.5% of rubric weight) discriminates most between models; empathy/support is nearly flat across models; older models score poorly on the reality-testing axis.

Methodology Notes

Synthetic conversation prefixes generated to match privacy-preserving summaries of real ChatGPT mental-health usage; rubrics written independently by two licensed clinicians per conversation and finalised by a third; criteria weighted from -10 to +10 and classified by an LLM into ten behavioural axes. Responses sampled four times per task at default API settings and graded item by item by GPT-5.6 Sol at high reasoning effort (an OpenAI model grading OpenAI and competitor models). Effectively single-turn (the model answers the final user turn); language comparisons are descriptive; user-rubric collection was limited to non-acute conversations; chat modality only. Dataset released as a zip with a canary string, with a request not to post examples online. Verified by fetching the paper PDF (HTTP 200, 32 pages, CreationDate 2026-09-23) and the dataset zip (HTTP 200, 2.1 MB); the announcement page was read via a Wayback raw capture because openai.com returns 403 to fetchers.

Authors

Ali Malik, Declan Grabb, Rebecca Soskin Hicks, Rahul K. Arora, Preston Bowman, Michael Sharman, Sahra Ghalebikesabi, Mikhail Trofimov, Vinnie Monaco, Karan Singhal

Tags

openaimentalhealthbenchbenchmarkrubricteen-personaexpert-cohortopen-datasetllm-judge

Cite This

APA

Ali Malik et al. (2026). MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations. OpenAI. https://cdn.openai.com/ctf-cdn/MentalHealthBench_A_Comprehensive_Benchmark_of_AI_Capabilities_in_Realistic_Mental_Health_Conversations.pdf