Making medical AI benchmarks clinically interpretable: the case of mental health
Argues that general medical AI benchmarks should report domain-specific results, and demonstrates this on HealthBench by isolating its mental health conversations. The authors compare mental health scores with overall scores across 16 model deployments and describe the topic composition and rubric coverage of the mental health subset.
Publisher
BMJ Mental Health (BMJ); RAND
Published
28 Sept 2026
Added
today
Key Findings
- A classifier identified 332 mental health conversations among HealthBench's 5,000; against a double-coded human sample (n=100) it had sensitivity 0.980, specificity 1.000 and F1 0.990.
- Across 16 Azure OpenAI deployments, mental health scores differed from overall HealthBench scores by -0.034 to +0.047.
- Topic composition was uneven: postpartum depression made up 26.8% of mental health conversations (n=89), against suicidal ideation 4.5% (n=15), psychosis 0.9% (n=3) and eating disorders 1.5% (n=5).
- Of 3,256 applicable rubric criteria, 13.2% explicitly assessed mental-health-specific behaviours such as risk recognition, escalation and non-reinforcement of harmful beliefs.
Methodology Notes
Perspective article with an original re-analysis of HealthBench using a published query-profile classifier (GPT-4o tagging) validated on a double-coded sample with third-reviewer adjudication; stratified sampling for model scoring. Open access (CC BY-NC 4.0); replication code released by the authors. Published online 2026-09-28 (Crossref and page citation metadata). Verified by fetching doi.org/10.1136/bmjment-2026-302921, which resolves to mentalhealth.bmj.com (HTTP 200), and the Crossref record.
Authors
Ryan K. McBain, Ellice Huang, Caroline Figueroa, Li Ang Zhang, Jonathan Cantor
Tags
Cite This
APA
Ryan K. McBain et al. (2026). Making medical AI benchmarks clinically interpretable: the case of mental health. BMJ Mental Health (BMJ); RAND. https://mentalhealth.bmj.com/content/29/1/e302921
Related Insights
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) · 25 Aug 2026
MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations
OpenAI · 23 Sept 2026