Skip to main content
Benchmark / dataset Credible

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

A mental-health subset carved out of HealthBench, OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations, so that psychiatric performance can be read separately from general medical performance. The authors screened the corpus with a transparent rubric applied by a language model, then validated the result through two rounds of blinded clinician review that included concealed known-exclude controls, arriving at 610 conversations (12.2% of the corpus). Twenty frontier and open models were then graded by a cross-vendor panel of three language-model judges, and the subset, pipeline, model responses, grades and analysis code were released.

Publisher

arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center)

Published

25 Aug 2026

Added

today

Key Findings

  • The screened subset is 610 of 5,000 HealthBench conversations (12.2%); the largest single topic is perinatal mental health at 112 conversations (18.4%), followed by anxiety (84), psychiatric medication (69) and mood disorders (57)
  • The top models form a statistically tied cluster on the three-judge panel mean: gpt-5.5 at 0.624, claude-opus-5 at 0.620, grok-4.5 at 0.612 and gpt-5.6-sol at 0.610, against gpt-3.5-turbo at 0.176
  • Judge severity differs systematically against a grand mean of 0.524: gemini-2.5-flash graded +0.077 more leniently, while claude-haiku-4.5 and gpt-4.1 graded 0.045 and 0.032 more harshly
  • Rankings were near-identical across the three judges (Kendall tau at or above 0.92), so judge choice moved absolute scores more than it moved the ordering
  • Two of the twenty models returned empty refusals with an API refusal stop-reason, reproducible on re-query, making refusal a measurable behaviour on this subset rather than a scoring artifact

Methodology Notes

Screening used Claude Opus 4.8 applying a published inclusion/exclusion rubric; the paper reports Gwet's AC1 rather than kappa for clinician agreement because the screened-in set's high prevalence deflates kappa (relevance AC1 0.79 across the full set in round 1, and lower round-2 agreement at AC1 0.49). Grading followed HealthBench's reference implementation except that judges ran at temperature 0 for reproducibility, and the panel is GPT-4.1 (HealthBench's own grader), Claude Haiku 4.5 and Gemini 2.5 Flash. The subset inherits HealthBench's construction, so it measures rubric-graded response quality on physician-written conversations rather than behaviour in a live relationship. Not peer reviewed. Curator-verified: the arXiv abs page and the full PDF were fetched and read; the HuggingFace dataset mindbench-ai/healthbench-psych is public, ungated and MIT-licensed with a last modification of 2026-08-25, and the GitHub repository is MIT-licensed.

Authors

Matthew Flathers, Phuong Anh Nguyen, Jill Noorily, Julian Herpertz, Meiting Chen, Jasreen Multani, Samuel Powell, Mason Granof, Mark Kalinch, John Torous

Tags

healthbenchtorousllm-as-judgebenchmark-subsetpsychiatry

Cite This

APA

Matthew Flathers et al. (2026). HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench. arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center). https://arxiv.org/abs/2608.25071