SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats
A preprint introducing a Chinese-language benchmark for contextual suicide-risk assessment in multi-party group chats, addressing the gap left by prior post-level social-media studies. Built from public group-chat data using signal-word extraction and bidirectional context expansion, with user risk levels annotated via an expert-validated, LLM-assisted paradigm. The authors benchmark pretrained language models and more than 40 LLMs, finding conversational context essential for reliable risk assessment and early detection in multi-party settings difficult.
Key Findings
- Contains 13,312 contextual segments from 1,406 users, drawn from 258,228 raw messages across 20 public Chinese suicide-related Telegram groups (Jan 2020-Jun 2025)
- Annotates fine-grained user risk levels via an expert-validated, LLM-assisted paradigm, spanning no risk through suicidal ideation, behaviour, and attempt
- Multi-message contextual information materially improves risk-assessment reliability over isolated post-level analysis
- Partial-context and fine-tuning experiments expose the difficulty of early suicide-risk detection in fragmented multi-party conversation
- The dataset is not publicly released due to ethical and sensitivity concerns; access is restricted to accredited mental-health and suicide-prevention research institutions on request
Methodology Notes
Public Telegram group chats; segments average roughly 19 messages / 535 Chinese characters / 4.6 participants. Expert-validated, LLM-assisted annotation of user risk levels. Evaluated pretrained language models plus 40+ LLMs. Chinese only; restricted-access dataset. arXiv preprint, not peer-reviewed; v1 submitted 27 May 2026.
Sources
arXiv abstract (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Xiangyu Wang, Zhiwei Yu, Chengze Du, Dingchang Wang, Yuhan Ye, Fangyu Zheng
Tags
Cite This
APA
Xiangyu Wang et al. (2026). SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats. arXiv. https://arxiv.org/abs/2605.27911
Related Insights
Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench)
arXiv (Chinese research team) · 2 Jun 2025
Expert-Level Crisis Detection in Mental Health Conversations
arXiv (Emory University-led) · 9 Jun 2026