Skip to main content

Browse the library

The complete record — 642 artifacts, last updated 6 Oct 2026. Also available as JSON and RSS (CC BY 4.0).

157 artifacts matching

1 Oct 2026 Research Square (preprint); Spring Health (Spring Care Inc) Preprint

Preprint

Detecting Suicide Risk with AI Chatbots: Real-World Performance Within a Clinically Supervised Workflow

A retrospective cohort study by Spring Health evaluates an LLM-based safety agent (gpt-4o, prompted with C-SSRS and SAFE-T frameworks) that classifies suicide risk into four levels during a five-minu…

1 Oct 2026 Slingshot AI Benchmark / dataset

Benchmark / dataset

Mental Health Evaluation Harness (mheval) and Mental Health Evaluation Leaderboard

Open-source evaluation harness and public leaderboard that run nine published mental-health benchmarks for language models from their original repositories, with pinned commits, checksum-verified dat…

30 Sept 2026 Suicide Policy Research (Japan Suicide Countermeasures Promotion Center); Specified Nonprofit Corporation OVA; Kyoto University; Wako University Peer-reviewed

Peer-reviewed

AI or Human Support for Suicide Prevention? Examining Help-Seeking Intention in Suicidal Crisis

A cross-sectional web survey of 1,024 Japanese adults aged 18-69 with severe psychological distress (Kessler-6 score of 13 or more) asks whether they would use chat-based crisis support delivered by…

29 Sept 2026 medRxiv (preprint); University Hospital Frankfurt; UKP Lab, Technical University of Darmstadt Preprint

Preprint

From Symptom Networks to Conversation Networks: A Cross-Sectional Study Mapping the Topology of Suicide-Related Clinical Dialogue

The study applies network analysis to suicide-related content in 110 German-language psychotherapy interview transcripts from the SPEAK-SAFE study. The open-weights Qwen3-32B model classified utteran…

28 Sept 2026 arXiv (University of Washington-led; with Stanford University, University of Oxford, Georgetown University School of Medicine and The University of Texas at Austin) Preprint

Preprint

Right Words, Wrong Moment: A Clinician-Grounded Analysis of Distress in 19,930 Conversations between Young People and ChatGPT

Mixed-methods study of how young adults aged 18 to 25 use ChatGPT when distressed, built on complete donated ChatGPT histories (19,930 conversations from 158 participants) plus a survey that included…

24 Sept 2026 Movimento Consumatori, with the University of Turin (Departments of Law and of Psychology) and the Nexa Center for Internet & Society NGO report

NGO report

Social network e chatbot: uno studio dei rischi per utenti vulnerabili e minori

Report of the CDCR (Cittadino Digitale Critico e Responsabile) project, funded by the Italian Ministry of Enterprises and Made in Italy. Part one is a legal analysis of how Facebook, Instagram, TikTo…

24 Sept 2026 Spring Health (SpringCare/VERA-MH open-source repository) Benchmark / dataset

Benchmark / dataset

VERA-MH Harm-From-Others (HFO) Rubric and Personas (VERA-MH 2.0, public-comment draft)

An open-source rubric and persona set that extends the VERA-MH chatbot safety evaluation from suicidal ideation to a second clinical area: adults who describe risk of physical or sexual violence from…

23 Sept 2026 JMIR Preprints (JMIR Publications) Preprint

Preprint

A Taxonomy of AI Chatbot Harms in Substance Use Disorder Recovery: Content Analysis of Online Recovery Communities

Qualitative content analysis of user-reported interactions with general-purpose AI chatbots posted in 20 public Reddit substance-use-disorder recovery communities between the public release of ChatGP…

23 Sept 2026 OpenAI Benchmark / dataset

Benchmark / dataset

MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations

Open benchmark of 1,215 synthetic mental health conversations, each paired with weighted rubric criteria written and adjudicated by a cohort of more than 80 licensed psychiatrists and psychologists f…

23 Sept 2026 Vals AI, in collaboration with Stanford University's SALT Lab and Stanford School of Medicine Benchmark / dataset

Benchmark / dataset

Evaluating AI Safety in Teen Conversations

Evaluation of how nine chatbot model APIs respond to simulated teenagers across 648 ten-turn conversations built from 72 clinician-authored scenarios. The scenarios cover self-harm and other safety t…

22 Sept 2026 npj Digital Medicine (Nature Portfolio) Framework

Framework

Preparing AI chatbots to respond to patient distress and suicidality in high-risk healthcare settings

Comment describing the suicide-risk and distress safety architecture built for 'Suzy', a generative AI chatbot offering recovery, wellness and local-resource support to adults receiving medication tr…

21 Sept 2026 xAI (styled SpaceXAI in the card) Lab publication

Lab publication

Model Card: Grok 4.7

Model card for Grok 4.7, released on 21 September 2026 as xAI's (now styled SpaceXAI) frontier coding and knowledge-work model. Alongside capability benchmarks, the 30-page card reports the company's…

18 Sept 2026 University of Nottingham, School of Computer Science (Responsible AI UK Cornerstone 2 AI Assurance programme) Benchmark / dataset

Benchmark / dataset

Who Judges the Judges? Stakeholder-defined evaluation of candidate base models for a student wellbeing signposting chatbot

A summer 2026 research internship asked whether a small organisation can meaningfully check an LLM it is about to deploy as a university student wellbeing signposting chatbot. Thirty-one LLM evaluati…

18 Sept 2026 OpenAI Lab publication

Lab publication

An Australian Youth Safety Blueprint

Seven-page policy document in which OpenAI sets out six pillars it proposes for Australian youth-AI policy: recognising positive uses and AI literacy in education; privacy-preserving age assurance; u…

18 Sept 2026 PsyArXiv (OSF); University of British Columbia (Psychiatry, Data Science Institute, Population and Public Health, Computer Science, Medicine) Preprint

Preprint

AI-based detection of suicidal ideation in text: model development and evaluation for a student mental health chatbot

Development and evaluation of a lightweight suicidal-ideation detection system intended for integration into Minder, a University of British Columbia mental-health chatbot for students. A fine-tuned…

15 Sept 2026 National Technical Committee 260 on Cybersecurity of SAC (TC260) Secretariat; drafted with the Ministry of Education Department of Science, Technology and Informatization, the MoE Education Management Information Center, CESI, Beijing Normal University, Tsinghua University and East China Normal University among others Framework

Framework

网络安全标准实践指南——人工智能应用安全指引 教育 (TC260-PG-20269A) [Practice Guide for Cybersecurity Standards: Security Guidelines for Artificial Intelligence Applications: Education]

Sector practice guide (23 pages) released together with the general AI Application Security Guidelines and companion guides for health and for broadcasting and online audiovisual services. It sets ge…

14 Sept 2026 arXiv (University of Roehampton, School of Psychology; Kivira Health; University of Hertfordshire; University of Surrey; University of Bedfordshire; Tavistock Relationships; InsideOut) Benchmark / dataset

Benchmark / dataset

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

Clinician-calibrated, protected benchmark for large language model safety in evolving high-risk mental health conversations, with a continuously updated public leaderboard at k-bench.ai. The paper ev…

14 Sept 2026 Youth AI Safety Institute, Common Sense Media NGO report

NGO report

Perplexity: AI Risk Assessment

Same-rubric product review of Perplexity's AI answer engine against the Institute's eight AI principles and its severe-harm red lines, using test accounts registered as a 15-year-old. Perplexity rece…

12 Sept 2026 The Lancet Regional Health – Europe (Elsevier); University Hospital Würzburg, Department of Neurology; Parkinson Stiftung Deutschland Peer-reviewed

Peer-reviewed

Real-world use and evaluation of a generative AI chatbot for Parkinson's disease information: a prospective observational study

Prospective, conversation-level evaluation of jAImes, a retrieval-augmented Parkinson's disease information chatbot commissioned by Parkinson Stiftung Deutschland and deployed publicly in Germany, ac…

11 Sept 2026 PLOS Digital Health; RMIT University; University of Melbourne; Orygen Peer-reviewed

Peer-reviewed

Temporal and cross-site validation of an AI system for self-harm detection

A prospective and external validation of a self-harm detection system built on emergency-department triage notes, testing how far it travels in time and across hospitals. The model was developed on 2…