Comparative analysis of BERT-based and generative large language models for detecting suicidal ideation: a performance evaluation study
A peer-reviewed study benchmarking three BERT variants (BERTimbau-Base, BERTimbau-Large, BERT-Multilingual) against three generative LLMs (ChatGPT-3.5, Bing/GPT-4, Bard) for detecting suicidal ideation in Brazilian Portuguese text. Using a psychologist-labelled corpus of 3,788 sentences with a held-out 100-sentence test set, it finds fine-tuned Portuguese-specific encoders competitive with, and cheaper to deploy than, large generative models, though a zero-shot generative model achieved the single best overall score.
Publisher
Cadernos de Saúde Pública
Published
25 Nov 2024
Added
2 weeks ago
Key Findings
- Zero-shot Bing/GPT-4 achieved approximately 98% across reported metrics, the best of any model tested
- Fine-tuned BERTimbau-Large reached about 96% accuracy and BERTimbau-Base about 94%, both outperforming zero-shot ChatGPT-3.5, Bard, and multilingual BERT
- Fine-tuning compact, language-specific encoder models was competitive with far larger generative models for this detection task
- Evaluation used non-clinical Brazilian Portuguese social-media text labelled by psychology professionals, not clinical records
Methodology Notes
Supervised fine-tuning of BERT variants with standard preprocessing and hold-out evaluation; generative LLMs evaluated zero-shot via prompt engineering. Dataset: 3,788 labelled sentences (2,691 negative, 1,097 positive) from Twitter/X — the Boamente corpus — with a balanced 100-sentence test subset. Published in Cadernos de Saúde Pública (Fiocruz), 2024;40(10):e00028824.
Sources
SciELO article (Cadernos de Saúde Pública) (primary)
Boamente Portuguese suicidal-ideation dataset (Zenodo)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Adonias Caetano de Oliveira, Renato Freitas Bessa, Ariel Soares Teles
Tags
Cite This
APA
Adonias Caetano de Oliveira, Renato Freitas Bessa, Ariel Soares Teles (2024). Comparative analysis of BERT-based and generative large language models for detecting suicidal ideation: a performance evaluation study. Cadernos de Saúde Pública. https://www.scielo.br/j/csp/a/XrbVfvybPj9tvJ8qWv7j8VC/?lang=en
Related Insights
Preliminary Evaluation of the Engagement and Effectiveness of a Mental Health Chatbot
Frontiers in Digital Health · 30 Nov 2020
Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench)
arXiv (Chinese research team) · 2 Jun 2025