Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement
Asks how well deployed safety signals recover clinically meaningful suicide-risk severity rather than a binary flag. Releases, under gated access, a benchmark of 516 r/SuicideWatch posts rated by a licensed psychiatrist on a four-level ordinal schema (Indicator, Ideation, Behavior, Attempt) grounded in the Columbia Suicide Severity Rating Scale, and evaluates vendor moderation APIs, prompted LLMs and supervised baselines under seven ordinal-aware metrics, framed by graded-response duties such as California Senate Bill 243.
Publisher
arXiv (Wondi AI; University of California, Berkeley; MIT; Harvard Medical School; McLean Hospital); accepted at the NLP for Positive Impact workshop, EMNLP 2026
Published
5 Sept 2026
Added
today
Key Findings
- OpenAI's omni-moderation endpoint separates low- from high-severity posts well (high-risk F1 0.860) but measures severity poorly (macro F1 0.395), systematically over-predicting the most severe category.
- Clinically grounded zero-shot prompting recovers much of the gap (macro F1 0.562); expert-authored clinical framing, rather than fine-tuning, added reasoning or naive multi-agent aggregation, is the effective lever.
- The value of reasoning depends on register: it hurts on long, noisy Reddit posts and helps on short, clinician-authored statements.
- Models compared under identical prompting include GPT-4o-mini, Gemini 2.5 Flash with safety filters disabled, Claude Haiku 4.5, GPT-4o and Claude Fable 5; Gemini's categorical safety ratings serve as a second API baseline; a 20-post three-psychiatrist reliability check reports pairwise quadratic-weighted kappa with wide intervals.
- The authors argue that a proportionate duty of care requires graded severity, not a binary flag, and release the evaluation framework.
Methodology Notes
Single primary psychiatrist annotator (Harvard Medical School and McLean Hospital affiliation) with a small three-rater reliability check on 20 posts; English Reddit data; benchmark gated for ethics reasons; first author affiliated with Wondi AI, a company in the suicide-risk measurement space (vendor interest). arXiv 2609.06263 version 1, 5 September 2026 (announced 9 September); 20 pages; accepted at NLP4PI, EMNLP 2026 (workshop paper).
Topics
Authors
Shreyas Krishnan, Gun Ahn, Jungjin Kim
Across NOPE's trackers
Tags
Cite This
APA
Shreyas Krishnan, Gun Ahn, Jungjin Kim. (2026). Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement. arXiv (Wondi AI; University of California, Berkeley; MIT; Harvard Medical School; McLean Hospital); accepted at the NLP for Positive Impact workshop, EMNLP 2026. https://arxiv.org/abs/2609.06263
Related Insights
Exploring generalizability and explainability of LLMs in classifying clinically rated suicidal ideation using heterogeneous data
npj Digital Medicine (Springer Nature) · 5 Sept 2026
Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench)
arXiv (Chinese research team) · 2 Jun 2025
Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment
Psychiatric Services (American Psychiatric Association); RAND-led author team · 26 Aug 2025
Ground Truths in Suicide Research: The Current State of AI-Based Suicide Detection in Social Media
Association for Computational Linguistics (Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology, CLPsych 2026) · 1 Jul 2026