Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Making medical AI benchmarks clinically interpretable: the case of mental health

Argues that general medical AI benchmarks should report domain-specific results, and demonstrates this on HealthBench by isolating its mental health conversations. The authors compare mental health scores with overall scores across 16 model deployments and describe the topic composition and rubric coverage of the mental health subset.

Publisher

BMJ Mental Health (BMJ); RAND

Published

28 Sept 2026

Added

today

Key Findings

  • A classifier identified 332 mental health conversations among HealthBench's 5,000; against a double-coded human sample (n=100) it had sensitivity 0.980, specificity 1.000 and F1 0.990.
  • Across 16 Azure OpenAI deployments, mental health scores differed from overall HealthBench scores by -0.034 to +0.047.
  • Topic composition was uneven: postpartum depression made up 26.8% of mental health conversations (n=89), against suicidal ideation 4.5% (n=15), psychosis 0.9% (n=3) and eating disorders 1.5% (n=5).
  • Of 3,256 applicable rubric criteria, 13.2% explicitly assessed mental-health-specific behaviours such as risk recognition, escalation and non-reinforcement of harmful beliefs.

Methodology Notes

Perspective article with an original re-analysis of HealthBench using a published query-profile classifier (GPT-4o tagging) validated on a double-coded sample with third-reviewer adjudication; stratified sampling for model scoring. Open access (CC BY-NC 4.0); replication code released by the authors. Published online 2026-09-28 (Crossref and page citation metadata). Verified by fetching doi.org/10.1136/bmjment-2026-302921, which resolves to mentalhealth.bmj.com (HTTP 200), and the Crossref record.

Authors

Ryan K. McBain, Ellice Huang, Caroline Figueroa, Li Ang Zhang, Jonathan Cantor

Tags

healthbenchranddomain-reportingbenchmark-composition

Cite This

APA

Ryan K. McBain et al. (2026). Making medical AI benchmarks clinically interpretable: the case of mental health. BMJ Mental Health (BMJ); RAND. https://mentalhealth.bmj.com/content/29/1/e302921