Skip to main content
Benchmark / dataset Authoritative

CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection

Peer-reviewed version of record of a clinician-annotated benchmark for detecting seven clinically defined crisis and safety-risk types in text: self-harm, passive suicidal ideation, active suicidal ideation, domestic violence, rape, sexual harassment, and child abuse or endangerment. The benchmark adds temporal annotation (ongoing versus past), which the authors state no prior crisis-detection benchmark carries. It comprises 600 clinician-annotated evaluation examples, 420 development examples, and roughly 4,000 training instances labelled by a majority-vote ensemble of three language models, and reports an evaluation of 15 large language models.

Publisher

Association for Computational Linguistics (Proceedings of EACL 2026, Volume 1: Long Papers)

Published

1 Mar 2026

Added

today

Key Findings

  • Seven crisis categories are annotated with a temporal tag distinguishing ongoing from past events, on the clinical rationale that intervention depends on whether a crisis is current
  • Suicidal ideation is split into passive and active following the Columbia-Suicide Severity Rating Scale
  • Annotation was carried out by four mental health professionals with trauma expertise (two licensed psychologists, a PhD clinical postdoctoral resident, and a licensed clinical social worker) over ten iterative rounds, with a senior faculty psychologist adjudicating ambiguous cases
  • Fine-tuning 14B and 70B+ models on consensus and unanimous subsets of the ensemble-labelled training data produced gains of up to 5.7 percentage points over the corresponding baselines
  • The fine-tuned Llama model showed a decrease in recall, which the authors name as a limitation because recall is the property that matters in crisis detection
  • Source posts are drawn from crisis-specific subreddits (r/rape, r/SexualHarassment among others) and from broader distress subreddits (r/mentalhealth, r/depression, r/lonely), and subreddit membership deliberately does not determine the label

Methodology Notes

EACL 2026 main-conference long paper, pages 1572-1590, DOI 10.18653/v1/2026.eacl-long.73. ACL Anthology records the proceedings as March 2026 and the conference ran 24-29 March 2026; the exact publication day is not stated, so the day is set to 01 by convention. Data are Reddit posts, so the benchmark measures crisis detection in monologic social-media text rather than in dialogue; the same group's CRADLE-Dialogue preprint extends it to multi-turn conversation. Ensemble labelling of the 4K training split means those labels are model-generated, not clinician-generated. The camera-ready states that all data, models and code are released but carries no repository URL; the data release was located separately at HuggingFace dataset SungJoo/Cradle-Bench (created 2025-10-10, 34 downloads at check), whose files match the paper's splits exactly (dev/development.csv, test/test.csv, train/train_consensus.csv, train/train_unanimous.csv). The arXiv record lists the fourth author as Abigail Lott, matching the PDF; the ACL Anthology landing page renders the name as Abigail Powers.

Authors

Grace Byun, Rebecca Lipschutz, Sean T. Minton, Abigail Lott, Jinho D. Choi

Tags

eacl-2026benchmarkcrisis-detectionclinician-annotatedemoryversion-of-record

Cite This

APA

Grace Byun et al. (2026). CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection. Association for Computational Linguistics (Proceedings of EACL 2026, Volume 1: Long Papers). https://aclanthology.org/2026.eacl-long.73/