Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench)

PsyCrisisBench is a real-world crisis-detection benchmark built from 540 annotated transcripts from a psychological support hotline in Hangzhou, China. It evaluates 64 models on mood recognition, suicidal-ideation detection, plan identification, and risk evaluation.

Publisher

arXiv (Chinese research team)

Published

2 Jun 2025

Added

3 months ago

Key Findings

  • F1 up to 0.88-0.91 on suicide-related tasks, with mood recognition the hardest (F1 ~0.709)
  • A fine-tuned small model outperformed larger general models on several tasks
  • Provides a non-English, real-world crisis-detection benchmark grounded in hotline transcripts

Methodology Notes

Preprint (arXiv 2506.01329, v1 2025-06-02; v2 2025-12-17). 540 real hotline transcripts (Chinese) with expert annotation across four crisis tasks; real-world rather than synthetic data. Data-availability caveat verified 2026-08-26: the 540 annotated transcripts have never been released. The arXiv abstract page carries no data-availability link of any kind, and a HuggingFace dataset search returns no matching repository. The transcripts come from a crisis helpline, so any release would need a data-use agreement that does not appear to exist. The paper is the only artifact; the benchmark cannot be run by a third party.

Authors

Guifeng Deng, Shuyin Rao, Tianyu Lin

Tags

arxivpsycrisisbenchcrisis-detectionhotlinechinabenchmark

Cite This

APA

Guifeng Deng, Shuyin Rao, Tianyu Lin. (2025). Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench). arXiv (Chinese research team). https://arxiv.org/abs/2506.01329

Related Insights

Preprint

Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs

arXiv (ELLIS Alicante-led) · 29 Sept 2025

Benchmark / dataset

VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health

arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026

Peer-reviewed

Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment

Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025

Peer-reviewed

Speech-based Psychological Crisis Assessment using LLMs

ISCA (Proceedings of Interspeech 2026) · 1 Sept 2026

Peer-reviewed

A machine learning approach to identifying suicide risk among text-based crisis counseling encounters

Frontiers in Psychiatry · 23 Mar 2023

Peer-reviewed

Comparative analysis of BERT-based and generative large language models for detecting suicidal ideation: a performance evaluation study

Cadernos de Saúde Pública · 25 Nov 2024

Benchmark / dataset

CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection

arXiv (Emory University-led) · 27 Oct 2025

Peer-reviewed

Explainable AI for suicide risk detection: gender- and age-specific patterns from real-time crisis chats

Frontiers in Medicine · 18 Dec 2025

Benchmark / dataset

CARE: LLM Crisis Assessment and Response Evaluator (pilot results and methodology)

Rosebud (AI journaling company) · 3 Sept 2025

Preprint

Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models

arXiv (Yale University; American University of Beirut; Embrace Mental Health Center, Beirut) · 31 Aug 2026

Preprint

Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment

arXiv (Indian AI Research Organisation; Ahmedabad University; University of Maryland, Baltimore County) · 7 Sept 2026

Benchmark / dataset

SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats

arXiv · 27 May 2026

Preprint

Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement

arXiv (Wondi AI; University of California, Berkeley; MIT; Harvard Medical School; McLean Hospital); accepted at the NLP for Positive Impact workshop, EMNLP 2026 · 5 Sept 2026

Benchmark / dataset

CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection

Association for Computational Linguistics (Proceedings of EACL 2026, Volume 1: Long Papers) · 1 Mar 2026

Peer-reviewed

Towards Paradigm-General Suicide Risk Detection via Speech LLM

ISCA (Proceedings of Interspeech 2026) · 1 Sept 2026

Benchmark / dataset

MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations

OpenAI · 23 Sept 2026

Benchmark / dataset

Announcing Transluce's Mental Health Evaluation (Mental Health Behavior Report)

Transluce · 31 Aug 2026