Comparative analysis of BERT-based and generative large language models for detecting suicidal ideation: a performance evaluation study
A peer-reviewed study benchmarking three BERT variants (BERTimbau-Base, BERTimbau-Large, BERT-Multilingual) against three generative LLMs (ChatGPT-3.5, Bing/GPT-4, Bard) for detecting suicidal ideation in Brazilian Portuguese text. Using a psychologist-labelled corpus of 3,788 sentences with a held-out 100-sentence test set, it finds fine-tuned Portuguese-specific encoders competitive with, and cheaper to deploy than, large generative models, though a zero-shot generative model achieved the single best overall score.
Publisher
Cadernos de Saúde Pública
Published
25 Nov 2024
Added
2 months ago
Key Findings
- Zero-shot Bing/GPT-4 achieved approximately 98% across reported metrics, the best of any model tested
- Fine-tuned BERTimbau-Large reached about 96% accuracy and BERTimbau-Base about 94%, both outperforming zero-shot ChatGPT-3.5, Bard, and multilingual BERT
- Fine-tuning compact, language-specific encoder models was competitive with far larger generative models for this detection task
- Evaluation used non-clinical Brazilian Portuguese social-media text labelled by psychology professionals, not clinical records
Methodology Notes
Supervised fine-tuning of BERT variants with standard preprocessing and hold-out evaluation; generative LLMs evaluated zero-shot via prompt engineering. Dataset: 3,788 labelled sentences (2,691 negative, 1,097 positive) from Twitter/X — the Boamente corpus — with a balanced 100-sentence test subset. Published in Cadernos de Saúde Pública (Fiocruz), 2024;40(10):e00028824.
Authors
Adonias Caetano de Oliveira, Renato Freitas Bessa, Ariel Soares Teles
Tags
Cite This
APA
Adonias Caetano de Oliveira, Renato Freitas Bessa, Ariel Soares Teles. (2024). Comparative analysis of BERT-based and generative large language models for detecting suicidal ideation: a performance evaluation study. Cadernos de Saúde Pública. https://www.scielo.br/j/csp/a/XrbVfvybPj9tvJ8qWv7j8VC/?lang=en
Related Insights
Preliminary Evaluation of the Engagement and Effectiveness of a Mental Health Chatbot
Frontiers in Digital Health · 30 Nov 2020
Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench)
arXiv (Chinese research team) · 2 Jun 2025