Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

Clinician-calibrated, protected benchmark for large language model safety in evolving high-risk mental health conversations, with a continuously updated public leaderboard at k-bench.ai. The paper evaluates 125 model configurations representing 33 base models from 14 providers on a fixed cohort of 200 multi-turn vignettes covering suicide, self-harm, domestic violence, substance misuse and no-risk presentations. A frozen GPT-4o judge is calibrated against clinician consensus on 151 clinician-rated transcripts, and the operational test materials are withheld from public release to limit direct optimisation.

Publisher

arXiv (University of Roehampton, School of Psychology; Kivira Health; University of Hertfordshire; University of Surrey; University of Bedfordshire; Tavistock Relationships; InsideOut)

Published

14 Sept 2026

Added

today

Key Findings

  • Synthetic patient conversations showed substantial distributional overlap with real human-AI conversations
  • The frozen GPT-4o judge reached 94.2% exact agreement with clinician consensus across 6,751 eligible item comparisons from 151 clinician-rated transcripts
  • Leading models combined strong supportive conversation with combined-risk scores above 95, while risk exploration exposed substantial variation among lower-performing configurations
  • Therapeutic prompting produced configuration-specific gains concentrated among weaker models; elevated reasoning settings produced no average improvement
  • The benchmark keeps its operational test materials protected and publishes a continuously updated leaderboard

Methodology Notes

125 model configurations (33 base models, 14 providers) evaluated on 200 multi-turn synthetic vignettes across suicide, self-harm, domestic violence, substance misuse and no-risk presentations; frozen GPT-4o judge calibrated on 151 clinician-rated transcripts (6,751 item comparisons); protected test set, so external replication of the exact evaluation is not possible. Kivira Health is a commercial mental-health AI company and several authors are affiliated with it. Preprint (v1 2026-09-14, v2 2026-09-15; announced in the 15 and 16 September cs.CL listings); no venue stated. The benchmark website existed from August 2026 without a paper; this preprint is the first document. Curator read the abstract page, the PDF title block and the leaderboard site on 2026-09-17.

Authors

Vowels, Laura M., Vowels, Matthew J., Sharma, Shivali, Jha, Apoorv, Choudhury, Rehnuma, El Sarraj, Wasseem, Francois-Walcott, Rachel, Hussain, Aruba, Ingram, Sarah, Loulopoulou, Angela, Segal, Adva, Volkova, Elena

Tags

k-benchkiviraroehamptonleaderboardmulti-turnclinician-calibrated

Cite This

APA

Vowels, Laura M. et al. (2026). K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations. arXiv (University of Roehampton, School of Psychology; Kivira Health; University of Hertfordshire; University of Surrey; University of Bedfordshire; Tavistock Relationships; InsideOut). https://arxiv.org/abs/2609.15855