Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Exploring generalizability and explainability of LLMs in classifying clinically rated suicidal ideation using heterogeneous data

Hong Kong study asking whether a language-model classifier of clinician-rated suicidal ideation performs unequally across patient subgroups because of linguistic heterogeneity. Cantonese clinical-interview transcripts from 299 adults (194 outpatients with major depressive disorder and 105 community controls) were classified with a BERT model against interviewer ratings on HDRS item 11; alexithymia (TAS-20) defined the subgroups. A topic-general classifier was compared with mood-specific and suicide-specific classifiers built by factorising transcripts by conversation topic.

Publisher

npj Digital Medicine (Springer Nature)

Published

5 Sept 2026

Added

today

Key Findings

  • The topic-general classifier reached AUC 0.88 in non-alexithymic participants but 0.77 in alexithymic participants, with sensitivity 0.40 versus 0.31; whole-sample AUC 0.83, sensitivity 0.31, specificity 0.94.
  • The odds of missing a suicidal case were 2.39 times higher for alexithymic individuals (OR 2.39, p = 0.002).
  • Topic factorisation cut the subgroup AUC gap from 0.11 to 0.05 (mood-related classifier) and 0.01 (suicide-specific classifier); the suicide-specific classifier reached whole-sample AUC 0.96 with sensitivity 0.87 and specificity 0.91.
  • Suicidal ideation was rated present in 68 of 299 participants; prevalence was 36.8% in the alexithymia group versus 4.6% in the non-alexithymia group.

Methodology Notes

Single-site convenience sample recruited October 2020 to May 2022; suicidal ideation rated by trained interviewers on HDRS item H11 (rating of 2 or more as the reference standard; inter-rater kappa 0.92); 85 participants missing TAS-20 data. The 'LLM' of the title is a Chinese BERT binary classifier; reporting follows the TRIPOD-LLM checklist. Interview transcripts, not chatbot conversations. Data available on request only; code stated at github.com/LongdiXian/Suicide_LLM_Generalizability. Competing interests: two authors report fees or travel support from Eisai, Lundbeck HK and Aculys Pharma. Published online 2026-09-05 (article in press), CC BY 4.0; verified from the publisher's reference PDF (8 pages).

Authors

Rong Huang, Longdi Xian, Christopher Chi Wai Cheng, Jie Chen, Kit Ying Chan, Calvin Lam, Joey W. Y. Chan, Steven W. H. Chau, Ngan Yin Chan, Bei Huang, Yun Kwok Wing, Tim M. H. Li

Tags

suicidal-ideationalexithymiaclassifier-equitycantoneseberthong-kong

Cite This

APA

Rong Huang et al. (2026). Exploring generalizability and explainability of LLMs in classifying clinically rated suicidal ideation using heterogeneous data. npj Digital Medicine (Springer Nature). https://www.nature.com/articles/s41746-026-03198-w