Skip to main content

Browse the library

The complete record — 359 artifacts, last updated 10 Sept 2026. Also available as JSON and RSS (CC BY 4.0).

63 artifacts matching

10 Sept 2026 Medicines and Healthcare products Regulatory Agency (MHRA), National Commission into the Regulation of AI in Healthcare Government report

Government report

National Commission into the Regulation of AI in Healthcare: Recommendations for a future regulatory framework

Final report of the UK National Commission into the Regulation of AI in Healthcare, established in September 2025 to advise government on a regulatory framework for software and AI-enabled health tec…

8 Sept 2026 arXiv (Grow Therapy; Stanford University School of Medicine) Preprint

Preprint

Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions

Deployment report from Grow Therapy, a US behavioural-health company whose network of more than 25,000 licensed clinicians offers clients an AI coaching tool for use between therapy sessions. Drawing…

8 Sept 2026 World Journal of Psychiatry (Baishideng Publishing Group) Peer-reviewed

Peer-reviewed

Feasibility of human-in-the-loop multimodal generative artificial intelligence chatbot for school-based adolescent mental health support

Prospective, non-randomized, waitlist-controlled pilot at a junior high school in Zhejiang Province, China, of 'Duoduo', an acceptance-and-commitment-therapy-informed multimodal generative-AI chatbot…

5 Sept 2026 arXiv (Wondi AI; University of California, Berkeley; MIT; Harvard Medical School; McLean Hospital); accepted at the NLP for Positive Impact workshop, EMNLP 2026 Preprint

Preprint

Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement

Asks how well deployed safety signals recover clinically meaningful suicide-risk severity rather than a binary flag. Releases, under gated access, a benchmark of 516 r/SuicideWatch posts rated by a l…

5 Sept 2026 npj Digital Medicine (Springer Nature) Peer-reviewed

Peer-reviewed

Exploring generalizability and explainability of LLMs in classifying clinically rated suicidal ideation using heterogeneous data

Hong Kong study asking whether a language-model classifier of clinician-rated suicidal ideation performs unequally across patient subgroups because of linguistic heterogeneity. Cantonese clinical-int…

4 Sept 2026 OMEGA - Journal of Death and Dying (Sage) Peer-reviewed

Peer-reviewed

Dual Illegibility, Ambiguous Loss, and Disenfranchised Grief: A Hybrid Model for Understanding AI Companion Bereavement

Conceptual paper proposing that the distress some users experience when an AI companion is altered, restricted or lost is best understood as a technologically mediated hybrid loss at the intersection…

3 Sept 2026 arXiv (National University of Singapore, Department of Computer Science) Benchmark / dataset

Benchmark / dataset

LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues

Synthetic benchmark suite for estimating depression severity and its change across multi-session counselling dialogue: three independently generated editions totalling 7,749 five-session trajectories…

1 Sept 2026 JMIR Mental Health (JMIR Publications) Peer-reviewed

Peer-reviewed

Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation

Evaluates how well a GPT-4o-based chatbot with a structured system prompt delivers a single-session behavioural-activation intervention for people with depression aged 14 to 29. Forty-eight sessions…

31 Aug 2026 arXiv (Yale University; American University of Beirut; Embrace Mental Health Center, Beirut) Preprint

Preprint

Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models

Evaluates suicide-risk classification from de-identified transcripts of calls to Lebanon's National Lifeline for Emotional Support and Suicide Prevention. Calls were transcribed on site with a Levant…

31 Aug 2026 Journal of Marital and Family Therapy (Wiley, for the American Association for Marriage and Family Therapy) Peer-reviewed

Peer-reviewed

Are You Qualified, ChatGPT? Examining Clinical Skills and Competencies of ChatGPT in Delivering Systemic Interventions

Pilot study assessing the clinical and cultural competence of a custom GPT configured to deliver systemic (couple and family therapy) interventions aligned with the AAMFT Code of Ethics and core comp…

31 Aug 2026 JMIR AI (JMIR Publications) Peer-reviewed

Peer-reviewed

Generative Large Language Models in Mental Health Care Settings: Systematic Review and Meta-Analysis

PRISMA systematic review of quantitative studies evaluating general-purpose large language models for direct mental health care tasks, searched across PubMed, Embase, ACM Digital Library, IEEE Xplore…

25 Aug 2026 arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) Benchmark / dataset

Benchmark / dataset

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

A mental-health subset carved out of HealthBench, OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations, so that psychiatric performance can be read separately from general me…

21 Aug 2026 arXiv (preprint) Preprint

Preprint

Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy

The authors introduce an ontology of ten therapeutic moves, compact function-based categories grounded in the MULTI-60 psychotherapy process inventory, validated through an annotation campaign with f…

19 Aug 2026 JMIR Mental Health (JMIR Publications) Peer-reviewed

Peer-reviewed

Generative AI in Youth Mental Health Apps: Rapid Review

A rapid review of how generative AI has been integrated into mental health apps for young people, and of what the evaluation literature reports about benefits and disadvantages. The authors searched…

18 Aug 2026 U.S. Food and Drug Administration, Center for Devices and Radiological Health Regulator study

Regulator study

Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback

A discussion paper from FDA's device centre seeking public comment on how generative-AI-enabled medical devices should be regulated. It proposes distinguishing informational functions from action-dir…

17 Aug 2026 Journal of Psychopathology and Clinical Science (American Psychological Association) Peer-reviewed

Peer-reviewed

A Framework for Evidence-Based Psychotherapy with AI (EBP-AI)

The authors propose EBP-AI, a named framework of eight principles for building clinical AI applications that produce durable change rather than momentary relief, paired with technical questions for d…

15 Jul 2026 Proceedings of the IASEAI Conference (published by AAAI) Peer-reviewed

Peer-reviewed

Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk

Conference paper from UL Research Institutes' Digital Safety group and Sentio University assessing whether a general-purpose open-weight model follows clinically derived guidelines for ethical commun…

9 Jul 2026 The Lancet Psychiatry (Elsevier) Peer-reviewed

Peer-reviewed

Multidisciplinary research priorities for artificial intelligence in mental health: a call to action

A Position Paper by 20 authors across Australia, the UK, the Netherlands, Brazil, Germany, China, Sweden, Czechia and the United States setting out a coordinated roadmap for the responsible evaluatio…

1 Jul 2026 Association for Computational Linguistics (Proceedings of ACL 2026, Volume 1: Long Papers) Peer-reviewed

Peer-reviewed

Responsible Evaluation of AI for Mental Health

Position-plus-analysis paper from a consortium of NLP and clinical-psychology researchers. Two annotators coded 135 mental-health papers published in ACL Anthology venues over five years (36% from 20…

1 Jul 2026 Association for Computational Linguistics (Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology, CLPsych 2026) Peer-reviewed

Peer-reviewed

The Reliability Illusion in Synthetic Patients: Psychometric Misalignment of Open-weight LLMs on PHQ-9 and GAD-7

Tests whether language models used as simulated patients produce psychometrically valid responses to the PHQ-9 and GAD-7 depression and anxiety scales. Four open-weight models (Llama-3.1-8B-Instruct,…