Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench)
PsyCrisisBench is a real-world crisis-detection benchmark built from 540 annotated transcripts from a psychological support hotline in Hangzhou, China. It evaluates 64 models on mood recognition, suicidal-ideation detection, plan identification, and risk evaluation.
Publisher
arXiv (Chinese research team)
Published
2 Jun 2025
Added
3 months ago
Key Findings
- F1 up to 0.88-0.91 on suicide-related tasks, with mood recognition the hardest (F1 ~0.709)
- A fine-tuned small model outperformed larger general models on several tasks
- Provides a non-English, real-world crisis-detection benchmark grounded in hotline transcripts
Methodology Notes
Preprint (arXiv 2506.01329, v1 2025-06-02; v2 2025-12-17). 540 real hotline transcripts (Chinese) with expert annotation across four crisis tasks; real-world rather than synthetic data. Data-availability caveat verified 2026-08-26: the 540 annotated transcripts have never been released. The arXiv abstract page carries no data-availability link of any kind, and a HuggingFace dataset search returns no matching repository. The transcripts come from a crisis helpline, so any release would need a data-use agreement that does not appear to exist. The paper is the only artifact; the benchmark cannot be run by a third party.
Sources
arXiv abstract(opens in a new tab) (primary)
Related: multi-label crisis detection on 1,057 Hangzhou hotline calls (PLOS Digital Health)(opens in a new tab) (13 May 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Guifeng Deng, Shuyin Rao, Tianyu Lin
Tags
Cite This
APA
Guifeng Deng, Shuyin Rao, Tianyu Lin. (2025). Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench). arXiv (Chinese research team). https://arxiv.org/abs/2506.01329
Related Insights
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
arXiv (ELLIS Alicante-led) · 29 Sept 2025
VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health
arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026
Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment
Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025
Speech-based Psychological Crisis Assessment using LLMs
ISCA (Proceedings of Interspeech 2026) · 1 Sept 2026
A machine learning approach to identifying suicide risk among text-based crisis counseling encounters
Frontiers in Psychiatry · 23 Mar 2023
Comparative analysis of BERT-based and generative large language models for detecting suicidal ideation: a performance evaluation study
Cadernos de Saúde Pública · 25 Nov 2024
CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection
arXiv (Emory University-led) · 27 Oct 2025
Explainable AI for suicide risk detection: gender- and age-specific patterns from real-time crisis chats
Frontiers in Medicine · 18 Dec 2025
CARE: LLM Crisis Assessment and Response Evaluator (pilot results and methodology)
Rosebud (AI journaling company) · 3 Sept 2025
Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models
arXiv (Yale University; American University of Beirut; Embrace Mental Health Center, Beirut) · 31 Aug 2026
Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment
arXiv (Indian AI Research Organisation; Ahmedabad University; University of Maryland, Baltimore County) · 7 Sept 2026
SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats
arXiv · 27 May 2026
Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement
arXiv (Wondi AI; University of California, Berkeley; MIT; Harvard Medical School; McLean Hospital); accepted at the NLP for Positive Impact workshop, EMNLP 2026 · 5 Sept 2026
CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection
Association for Computational Linguistics (Proceedings of EACL 2026, Volume 1: Long Papers) · 1 Mar 2026
Towards Paradigm-General Suicide Risk Detection via Speech LLM
ISCA (Proceedings of Interspeech 2026) · 1 Sept 2026
MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations
OpenAI · 23 Sept 2026
Announcing Transluce's Mental Health Evaluation (Mental Health Behavior Report)
Transluce · 31 Aug 2026