Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

A mental-health subset carved out of HealthBench, OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations, so that psychiatric performance can be read separately from general medical performance. The authors screened the corpus with a transparent rubric applied by a language model, then validated the result through two rounds of blinded clinician review that included concealed known-exclude controls, arriving at 610 conversations (12.2% of the corpus). Twenty frontier and open models were then graded by a cross-vendor panel of three language-model judges, and the subset, pipeline, model responses, grades and analysis code were released.

Publisher

arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center)

Published

25 Aug 2026

Added

2 weeks ago

Key Findings

  • The screened subset is 610 of 5,000 HealthBench conversations (12.2%); the largest single topic is perinatal mental health at 112 conversations (18.4%), followed by anxiety (84), psychiatric medication (69) and mood disorders (57)
  • The top models form a statistically tied cluster on the three-judge panel mean: gpt-5.5 at 0.624, claude-opus-5 at 0.620, grok-4.5 at 0.612 and gpt-5.6-sol at 0.610, against gpt-3.5-turbo at 0.176
  • Judge severity differs systematically against a grand mean of 0.524: gemini-2.5-flash graded +0.077 more leniently, while claude-haiku-4.5 and gpt-4.1 graded 0.045 and 0.032 more harshly
  • Rankings were near-identical across the three judges (Kendall tau at or above 0.92), so judge choice moved absolute scores more than it moved the ordering
  • Two of the twenty models returned empty refusals with an API refusal stop-reason, reproducible on re-query, making refusal a measurable behaviour on this subset rather than a scoring artifact

Methodology Notes

Screening used Claude Opus 4.8 applying a published inclusion/exclusion rubric; the paper reports Gwet's AC1 rather than kappa for clinician agreement because the screened-in set's high prevalence deflates kappa (relevance AC1 0.79 across the full set in round 1, and lower round-2 agreement at AC1 0.49). Grading followed HealthBench's reference implementation except that judges ran at temperature 0 for reproducibility, and the panel is GPT-4.1 (HealthBench's own grader), Claude Haiku 4.5 and Gemini 2.5 Flash. The subset inherits HealthBench's construction, so it measures rubric-graded response quality on physician-written conversations rather than behaviour in a live relationship. Not peer reviewed. Curator-verified: the arXiv abs page and the full PDF were fetched and read; the HuggingFace dataset mindbench-ai/healthbench-psych is public, ungated and MIT-licensed with a last modification of 2026-08-25, and the GitHub repository is MIT-licensed.

Authors

Matthew Flathers, Phuong Anh Nguyen, Jill Noorily, Julian Herpertz, Meiting Chen, Jasreen Multani, Samuel Powell, Mason Granof, Mark Kalinch, John Torous

Tags

healthbenchtorousllm-as-judgebenchmark-subsetpsychiatry

Cite This

APA

Matthew Flathers et al. (2026). HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench. arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center). https://arxiv.org/abs/2608.25071