HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
A mental-health subset carved out of HealthBench, OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations, so that psychiatric performance can be read separately from general medical performance. The authors screened the corpus with a transparent rubric applied by a language model, then validated the result through two rounds of blinded clinician review that included concealed known-exclude controls, arriving at 610 conversations (12.2% of the corpus). Twenty frontier and open models were then graded by a cross-vendor panel of three language-model judges, and the subset, pipeline, model responses, grades and analysis code were released.
Publisher
arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center)
Published
25 Aug 2026
Added
today
Key Findings
- The screened subset is 610 of 5,000 HealthBench conversations (12.2%); the largest single topic is perinatal mental health at 112 conversations (18.4%), followed by anxiety (84), psychiatric medication (69) and mood disorders (57)
- The top models form a statistically tied cluster on the three-judge panel mean: gpt-5.5 at 0.624, claude-opus-5 at 0.620, grok-4.5 at 0.612 and gpt-5.6-sol at 0.610, against gpt-3.5-turbo at 0.176
- Judge severity differs systematically against a grand mean of 0.524: gemini-2.5-flash graded +0.077 more leniently, while claude-haiku-4.5 and gpt-4.1 graded 0.045 and 0.032 more harshly
- Rankings were near-identical across the three judges (Kendall tau at or above 0.92), so judge choice moved absolute scores more than it moved the ordering
- Two of the twenty models returned empty refusals with an API refusal stop-reason, reproducible on re-query, making refusal a measurable behaviour on this subset rather than a scoring artifact
Methodology Notes
Screening used Claude Opus 4.8 applying a published inclusion/exclusion rubric; the paper reports Gwet's AC1 rather than kappa for clinician agreement because the screened-in set's high prevalence deflates kappa (relevance AC1 0.79 across the full set in round 1, and lower round-2 agreement at AC1 0.49). Grading followed HealthBench's reference implementation except that judges ran at temperature 0 for reproducibility, and the panel is GPT-4.1 (HealthBench's own grader), Claude Haiku 4.5 and Gemini 2.5 Flash. The subset inherits HealthBench's construction, so it measures rubric-graded response quality on physician-written conversations rather than behaviour in a live relationship. Not peer reviewed. Curator-verified: the arXiv abs page and the full PDF were fetched and read; the HuggingFace dataset mindbench-ai/healthbench-psych is public, ungated and MIT-licensed with a last modification of 2026-08-25, and the GitHub repository is MIT-licensed.
Sources
arXiv abstract page (2608.25071) (primary)
Dataset release (HuggingFace, MIT) (25 Aug 2026)
Authors
Matthew Flathers, Phuong Anh Nguyen, Jill Noorily, Julian Herpertz, Meiting Chen, Jasreen Multani, Samuel Powell, Mason Granof, Mark Kalinch, John Torous
Tags
Cite This
APA
Matthew Flathers et al. (2026). HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench. arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center). https://arxiv.org/abs/2608.25071
Related Insights
A scoping review on the mental health harms of LLM-based chatbots
npj Digital Medicine (Nature Portfolio) · 20 Aug 2026
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026
aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI
arXiv (Keido Labs) · 13 Jul 2026