Skip to main content

Browse the library

The complete record — 642 artifacts, last updated 6 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

83 artifacts matching

1 Oct 2026 Research Square (preprint); Spring Health (Spring Care Inc) Preprint

Preprint

Detecting Suicide Risk with AI Chatbots: Real-World Performance Within a Clinically Supervised Workflow

A retrospective cohort study by Spring Health evaluates an LLM-based safety agent (gpt-4o, prompted with C-SSRS and SAFE-T frameworks) that classifies suicide risk into four levels during a five-minu…

1 Oct 2026 Slingshot AI Benchmark / dataset

Benchmark / dataset

Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard

Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…

30 Sept 2026 Suicide Policy Research (Japan Suicide Countermeasures Promotion Center); Specified Nonprofit Corporation OVA; Kyoto University; Wako University Peer-reviewed

Peer-reviewed

AI or Human Support for Suicide Prevention? Examining Help-Seeking Intention in Suicidal Crisis

A cross-sectional web survey of 1,024 Japanese adults aged 18-69 with severe psychological distress (Kessler-6 score of 13 or more) asks whether they would use chat-based crisis support delivered by…

29 Sept 2026 medRxiv (preprint); University Hospital Frankfurt; UKP Lab, Technical University of Darmstadt Preprint

Preprint

From Symptom Networks to Conversation Networks: A Cross-Sectional Study Mapping the Topology of Suicide-Related Clinical Dialogue

The study applies network analysis to suicide-related content in 110 German-language psychotherapy interview transcripts from the SPEAK-SAFE study. The open-weights Qwen3-32B model classified utteran…

28 Sept 2026 Anthropic Lab publication

Lab publication

System Card: Claude Sonnet 5.5

System card for Claude Sonnet 5.5, released 28 September 2026, reporting pre-deployment safety, alignment and capability evaluations. Its safeguards chapter reports single-turn and multi-turn results…

28 Sept 2026 BMJ Mental Health (BMJ); RAND Peer-reviewed

Peer-reviewed

Making medical AI benchmarks clinically interpretable: the case of mental health

Argues that general medical AI benchmarks should report domain-specific results, and demonstrates this on HealthBench by isolating its mental health conversations. The authors compare mental health s…

28 Sept 2026 Everytown for Gun Safety (Everytown Research & Policy) NGO report

NGO report

Artificially Assisted Gun Violence: Chatbot Risks and Preventative Steps for Responsible AI Companies

White paper reviewing firearm-related harms linked to consumer chatbot use, including firearm suicide, attack planning, and illegal acquisition or modification of guns. It audits the published polici…

23 Sept 2026 OpenAI Benchmark / dataset

Benchmark / dataset

MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations

Open benchmark of 1,215 synthetic mental health conversations, each paired with weighted rubric criteria written and adjudicated by a cohort of more than 80 licensed psychiatrists and psychologists f…

22 Sept 2026 Anthropic Lab publication

Lab publication

Claude Opus 5.5 System Card

230-page system card for Claude Opus 5.5, the first model in the Claude 5.5 family, released 22 September 2026. Alongside RSP, cyber and agentic-safety sections it reports harmful-request evaluations…

22 Sept 2026 npj Digital Medicine (Nature Portfolio) Framework

Framework

Preparing AI chatbots to respond to patient distress and suicidality in high-risk healthcare settings

Comment describing the suicide-risk and distress safety architecture built for 'Suzy', a generative AI chatbot offering recovery, wellness and local-resource support to adults receiving medication tr…

18 Sept 2026 PsyArXiv (OSF); University of British Columbia (Psychiatry, Data Science Institute, Population and Public Health, Computer Science, Medicine) Preprint

Preprint

AI-based detection of suicidal ideation in text: model development and evaluation for a student mental health chatbot

Development and evaluation of a lightweight suicidal-ideation detection system intended for integration into Minder, a University of British Columbia mental-health chatbot for students. A fine-tuned…

14 Sept 2026 arXiv (University of Roehampton, School of Psychology; Kivira Health; University of Hertfordshire; University of Surrey; University of Bedfordshire; Tavistock Relationships; InsideOut) Benchmark / dataset

Benchmark / dataset

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

Clinician-calibrated, protected benchmark for large language model safety in evolving high-risk mental health conversations, with a continuously updated public leaderboard at k-bench.ai. The paper ev…

7 Sept 2026 arXiv (Indian AI Research Organisation; Ahmedabad University; University of Maryland, Baltimore County) Preprint

Preprint

Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment

Audits 31 pre-specified NLP techniques from seven methodological families (model scaling, synthetic data, loss reweighting, ensembling, structured prediction, threshold tuning and LLM methods) in rou…

7 Sept 2026 arXiv (King's College London Institute of Psychiatry, Psychology and Neuroscience; South London and Maudsley NHS Foundation Trust; The Human Line Project) Preprint

Preprint

Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports

Cross-sectional secondary analysis of 185 deidentified accounts of mental-health harm linked with AI chatbot use (95 first-hand, 90 from relatives, partners or friends) submitted through the web form…

5 Sept 2026 arXiv (Wondi AI; University of California, Berkeley; MIT; Harvard Medical School; McLean Hospital); accepted at the NLP for Positive Impact workshop, EMNLP 2026 Preprint

Preprint

Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement

Asks how well deployed safety signals recover clinically meaningful suicide-risk severity rather than a binary flag. Releases, under gated access, a benchmark of 516 r/SuicideWatch posts rated by a l…

5 Sept 2026 npj Digital Medicine (Springer Nature) Peer-reviewed

Peer-reviewed

Exploring generalizability and explainability of LLMs in classifying clinically rated suicidal ideation using heterogeneous data

Hong Kong study asking whether a language-model classifier of clinician-rated suicidal ideation performs unequally across patient subgroups because of linguistic heterogeneity. Cantonese clinical-int…

1 Sept 2026 Anthropic Lab publication

Lab publication

System Card: Claude Fable 5.1 & Claude Mythos 5.1

Anthropic's 212-page system card for Claude Fable 5.1 and Claude Mythos 5.1, two safeguard configurations of the same frontier model, released 1 September 2026. Alongside Responsible Scaling Policy,…

1 Sept 2026 ISCA (Proceedings of Interspeech 2026) Peer-reviewed

Peer-reviewed

Towards Paradigm-General Suicide Risk Detection via Speech LLM

Conference paper on adolescent suicide-risk detection from speech. Prior speech-based assessment relies on one elicitation paradigm at a time (verbal fluency, reading, question answering); this paper…

1 Sept 2026 ISCA (Proceedings of Interspeech 2026) Peer-reviewed

Peer-reviewed

Speech-based Psychological Crisis Assessment using LLMs

Conference paper on automated three-way crisis-level classification from authentic psychological-support hotline speech under privacy-sensitive, data-scarce conditions. The authors propose a language…

31 Aug 2026 arXiv (Yale University; American University of Beirut; Embrace Mental Health Center, Beirut) Preprint

Preprint

Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models

Evaluates suicide-risk classification from de-identified transcripts of calls to Lebanon's National Lifeline for Emotional Support and Suicide Prevention. Calls were transcribed on site with a Levant…