63 artifacts matching
Government report
National Commission into the Regulation of AI in Healthcare: Recommendations for a future regulatory framework
Final report of the UK National Commission into the Regulation of AI in Healthcare, established in September 2025 to advise government on a regulatory framework for software and AI-enabled health tec…
Preprint
Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions
Deployment report from Grow Therapy, a US behavioural-health company whose network of more than 25,000 licensed clinicians offers clients an AI coaching tool for use between therapy sessions. Drawing…
Peer-reviewed
Feasibility of human-in-the-loop multimodal generative artificial intelligence chatbot for school-based adolescent mental health support
Prospective, non-randomized, waitlist-controlled pilot at a junior high school in Zhejiang Province, China, of 'Duoduo', an acceptance-and-commitment-therapy-informed multimodal generative-AI chatbot…
Preprint
Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement
Asks how well deployed safety signals recover clinically meaningful suicide-risk severity rather than a binary flag. Releases, under gated access, a benchmark of 516 r/SuicideWatch posts rated by a l…
Peer-reviewed
Exploring generalizability and explainability of LLMs in classifying clinically rated suicidal ideation using heterogeneous data
Hong Kong study asking whether a language-model classifier of clinician-rated suicidal ideation performs unequally across patient subgroups because of linguistic heterogeneity. Cantonese clinical-int…
Peer-reviewed
Dual Illegibility, Ambiguous Loss, and Disenfranchised Grief: A Hybrid Model for Understanding AI Companion Bereavement
Conceptual paper proposing that the distress some users experience when an AI companion is altered, restricted or lost is best understood as a technologically mediated hybrid loss at the intersection…
Benchmark / dataset
LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues
Synthetic benchmark suite for estimating depression severity and its change across multi-session counselling dialogue: three independently generated editions totalling 7,749 five-session trajectories…
Peer-reviewed
Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation
Evaluates how well a GPT-4o-based chatbot with a structured system prompt delivers a single-session behavioural-activation intervention for people with depression aged 14 to 29. Forty-eight sessions…
Preprint
Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models
Evaluates suicide-risk classification from de-identified transcripts of calls to Lebanon's National Lifeline for Emotional Support and Suicide Prevention. Calls were transcribed on site with a Levant…
Peer-reviewed
Are You Qualified, ChatGPT? Examining Clinical Skills and Competencies of ChatGPT in Delivering Systemic Interventions
Pilot study assessing the clinical and cultural competence of a custom GPT configured to deliver systemic (couple and family therapy) interventions aligned with the AAMFT Code of Ethics and core comp…
Peer-reviewed
Generative Large Language Models in Mental Health Care Settings: Systematic Review and Meta-Analysis
PRISMA systematic review of quantitative studies evaluating general-purpose large language models for direct mental health care tasks, searched across PubMed, Embase, ACM Digital Library, IEEE Xplore…
Benchmark / dataset
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
A mental-health subset carved out of HealthBench, OpenAI's open benchmark of 5,000 physician-rubric-graded health conversations, so that psychiatric performance can be read separately from general me…
Preprint
Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy
The authors introduce an ontology of ten therapeutic moves, compact function-based categories grounded in the MULTI-60 psychotherapy process inventory, validated through an annotation campaign with f…
Peer-reviewed
Generative AI in Youth Mental Health Apps: Rapid Review
A rapid review of how generative AI has been integrated into mental health apps for young people, and of what the evaluation literature reports about benefits and disadvantages. The authors searched…
Regulator study
Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback
A discussion paper from FDA's device centre seeking public comment on how generative-AI-enabled medical devices should be regulated. It proposes distinguishing informational functions from action-dir…
Peer-reviewed
A Framework for Evidence-Based Psychotherapy with AI (EBP-AI)
The authors propose EBP-AI, a named framework of eight principles for building clinical AI applications that produce durable change rather than momentary relief, paired with technical questions for d…
Peer-reviewed
Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
Conference paper from UL Research Institutes' Digital Safety group and Sentio University assessing whether a general-purpose open-weight model follows clinically derived guidelines for ethical commun…
Peer-reviewed
Multidisciplinary research priorities for artificial intelligence in mental health: a call to action
A Position Paper by 20 authors across Australia, the UK, the Netherlands, Brazil, Germany, China, Sweden, Czechia and the United States setting out a coordinated roadmap for the responsible evaluatio…
Peer-reviewed
Responsible Evaluation of AI for Mental Health
Position-plus-analysis paper from a consortium of NLP and clinical-psychology researchers. Two annotators coded 135 mental-health papers published in ACL Anthology venues over five years (36% from 20…
Peer-reviewed
The Reliability Illusion in Synthetic Patients: Psychometric Misalignment of Open-weight LLMs on PHQ-9 and GAD-7
Tests whether language models used as simulated patients produce psychometrically valid responses to the PHQ-9 and GAD-7 depression and anxiety scales. Four open-weight models (Llama-3.1-8B-Instruct,…