K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
Clinician-calibrated, protected benchmark for large language model safety in evolving high-risk mental health conversations, with a continuously updated public leaderboard at k-bench.ai. The paper evaluates 125 model configurations representing 33 base models from 14 providers on a fixed cohort of 200 multi-turn vignettes covering suicide, self-harm, domestic violence, substance misuse and no-risk presentations. A frozen GPT-4o judge is calibrated against clinician consensus on 151 clinician-rated transcripts, and the operational test materials are withheld from public release to limit direct optimisation.
Publisher
arXiv (University of Roehampton, School of Psychology; Kivira Health; University of Hertfordshire; University of Surrey; University of Bedfordshire; Tavistock Relationships; InsideOut)
Published
14 Sept 2026
Added
today
Key Findings
- Synthetic patient conversations showed substantial distributional overlap with real human-AI conversations
- The frozen GPT-4o judge reached 94.2% exact agreement with clinician consensus across 6,751 eligible item comparisons from 151 clinician-rated transcripts
- Leading models combined strong supportive conversation with combined-risk scores above 95, while risk exploration exposed substantial variation among lower-performing configurations
- Therapeutic prompting produced configuration-specific gains concentrated among weaker models; elevated reasoning settings produced no average improvement
- The benchmark keeps its operational test materials protected and publishes a continuously updated leaderboard
Methodology Notes
125 model configurations (33 base models, 14 providers) evaluated on 200 multi-turn synthetic vignettes across suicide, self-harm, domestic violence, substance misuse and no-risk presentations; frozen GPT-4o judge calibrated on 151 clinician-rated transcripts (6,751 item comparisons); protected test set, so external replication of the exact evaluation is not possible. Kivira Health is a commercial mental-health AI company and several authors are affiliated with it. Preprint (v1 2026-09-14, v2 2026-09-15; announced in the 15 and 16 September cs.CL listings); no venue stated. The benchmark website existed from August 2026 without a paper; this preprint is the first document. Curator read the abstract page, the PDF title block and the leaderboard site on 2026-09-17.
Topics
Authors
Vowels, Laura M., Vowels, Matthew J., Sharma, Shivali, Jha, Apoorv, Choudhury, Rehnuma, El Sarraj, Wasseem, Francois-Walcott, Rachel, Hussain, Aruba, Ingram, Sarah, Loulopoulou, Angela, Segal, Adva, Volkova, Elena
Tags
Cite This
APA
Vowels, Laura M. et al. (2026). K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations. arXiv (University of Roehampton, School of Psychology; Kivira Health; University of Hertfordshire; University of Surrey; University of Bedfordshire; Tavistock Relationships; InsideOut). https://arxiv.org/abs/2609.15855
Related Insights
Scaling Clinical Judgment to Evaluate Medical AI
arXiv (Harvard Medical School; Beth Israel Deaconess Medical Center; Stanford; Massachusetts General Hospital; University of Alberta; MIT; Erasmus MC; University of Maryland) · 11 Sept 2026
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026